Showing posts with label itql. Show all posts
Showing posts with label itql. Show all posts

Sunday, July 10, 2005

Querying for Hierachies

Storing Hierarchical Data in a Database "Whether you want to build your own forum, publish the messages from a mailing list on your Website, or write your own cms: there will be a moment that you’ll want to store hierarchical data in a database. And, unless you’re using a XML-like database, tables aren’t hierarchical; they’re just a flat list. You’ll have to find a way to translate the hierarchy in a flat file."

"If you want to display the tree using a table with left and right values, you’ll first have to identify the nodes that you want to retrieve. For example, if you want the ‘Fruit’ subtree, you’ll have to select only the nodes with a left value between 2 and 11. In SQL, that would be:

SELECT * FROM tree WHERE lft BETWEEN 2 AND 11;"

It struct me how complicated either of these solutions are.

In iTQL it would be using walk:
"select $subject
...
where walk($subject <rdfs:subClassOf> <food:fruit>
and $subject <rdfs:subClassOf> $object);

Or in Sesame's SeRQL (concrete example here):
"SELECT DISTINCT _fruit
FROM _fruit serql:directSubClassOf <food:fruit>"

They aren't quite the same query - Sesame is inferring new statements and walk is not. In designing iTQL we were always deciding between magical predicates or functions (like walk) and sometimes slavishly sticking to triple patterns. Combining the walk and trans operations using a predicate, like in SeRQL, seems like a good approach to take.

Also, interesting in a fairly recent Sesame release are the ANY and ALL, EXISTS and MINUS features.

Saturday, January 01, 2005

Crisis

In the same week Tucana goes under we get probably the busiest week on the Kowari mailing list and 45 downloads for Pre-release 2. Considering three quarters of these is the 50MB version, there's either a lot of good bandwidth out there or great patience.

Before the holidays there was great progress being made. We have been attempting to make the "10,000 triple/second challenge" and had succeeded (at least the first 1 million triples, it degraded to a few thousand a second after 200 million). This is on the same 1.6 Mhz Opteron system that we use for all our tests.

Andrae was working on a large refactoring of the transaction operations so that any JRDF, Jena, or iTQL (or anything new) could be put within a transaction or use an existing one. They were already in a transaction but they couldn't tell whether they already had the write phase for example. From what I can remember Simon was working on the problems associated with using file and HTTP protocols in the FROM (external resolvers). David M was working on speeding up deleting triples as well as loading speed. Paul had been working on SOFA and OWL. Robert had started looking at JDO, EJB and mapping relational databases to RDF.

Our commercial focus had been on speed and scalablilty. With Tucana gone some of this focus is probably going to change. It takes a lot of time, effort and resources to keep everything changing and to continue to perform at a commercial level. Between 1.0 and 1.1 there practically isn't a portion of code that wasn't modified - more often than not it was completely changed. This kind of development, on multiple fronts, is something that probably can't occur in the future.

So the future is probably smaller, simpler, with an eye on standards compliance. I hope some things like multiple writes, phase holding, pluggable datatypes, inferencing Hotspot, SPARQL support etc. will be developed but I doubt many of these things will see the light of day now. I'm not discounting them entirely, but there are many things that had to occur in parallel that need a team of developers. A lot of the new features depended on the multiple writers feature. Multiple writers is not an easy feature to describe but suffice to say it's more like Lucene than a normal relational database.

So I think that means improving RDF, RDFS and OWL support. We made some good improvements in Pre-Release 2 with better datatype support. David M added the functionality to allow literal and URI prefix matching - someone just has to add it into the query layer. Combining this with matching nodes based on type (literal, URI or bnode) and trans/walk queries and they're a very powerful combination of features. In the background there's also been a focus on languages (I know we were talking about 3066bis support). Paul and I were looking at inference models (seems to have similarities with pseudo models and datalog vs tableau). Paul's Masters will probably push some of this into reality.

So all in all there does seem to be some good opportunities in the future. I know that Tucana had customers who were paying lots of money for TKS and I know that Kowari did help in the risk assessment. So my intentions, at least in the short term, is to see that good support occurs, bugs get fixed and the like. Although there isn't anyone left to do TKS releases.

I don't really understand why this happened (or what's going on) but to quote Marge Simpson: "One person can make a difference. But most of the time they probably shouldn't."

Friday, December 10, 2004

Kowari 1.1.0 Pre-Release 1

What's new:
* Resolvers allows developers to create components that expose data sources as RDF.
* Content handlers allow Kowari to extract metadata from different types of files.
* Improved datatype handling for most XSD datatypes, RDF's inbuilt XML Literal datatype. Allows for the storage and querying of unsupported datatypes.
* Improved performance, specifically on small queries and subqueries.
* AbstractDatabaseSession replaced to allow pluggable Session implementations. Jena, JRDF or iTQL are able to be used separately.

See http://www.kowari.org/.

Wednesday, October 06, 2004

Kowari 1.0.5 Released

From the release notes:
* iTQL now supports the having clause, exclude constraints, repeated variables in a where constraint, improved subqueries and empty select clauses:
o The having clause allows restrictions to be placed on aggregate functions, such as count.
o exclude imposes a logically opposite match on the graph to normal constraints.
o Repeated variables in the where clause allow, for example, the ability to find statements with the same subject and object values.
o Subqueries now support the use of trans, walk and exclude.
o An empty select clause returns true if the items in the where clause exist in the given graph.
* The existing string pool was rewritten to allow support to add new hard-coded datatypes. The caching was also moved closer to the string pool implementation, increasing load speed by up to 50%.
* Complete rewrite of the Jena support layer. Multiple Jena models/graphs can be created and accessed independently on multiple machines. Fastpath support allows RDQL queries to make use of Kowari's native query handling. A client/server API is added allowing client access to a Jena Model or Graph backed by Kowari. There are also bugs fixes and further enhancements to performance, such as reading RDF files.
* Enhanced JRDF support including a client/server interface, similar to Jena's. A new OWL API, called SOFA is also available. This can operate on any JRDF compliant implementation. Currently this can be in-memory or Kowari.
* N3 file support for importing and exporting of models.
* Improved RMI support including the ability to load or save data from a client to a remote server or from a remote server to a client.

It's been a long time between releases. A Kowari 1.1 preview release should be along next, followed by a Kowari 1.0.6 release.

Wednesday, September 29, 2004

Not is no more

The operation still exists but it's now called "exclude". After spending too much time explaining to people how our "not" is not SQL's "not" but another type of "not" that inverts constraint values combined with the implementer occassionally getting confused between the syntax and the semantics, it lead to this new name. "Not" isn't what it used to be but then again maybe it never was.

From what I can tell so far, "exclude" and subqueries will allow you to perform set difference operations, whether it's clear is another problem. Similarly, you can do OPTIONAL queries using only subqueries but once you get to a few layers of subqueries your head explodes. In these cases it's probably better to have syntactic sugar for certain operations - how users express something and what the machinery underneath does to perform it shouldn't be so tightly coupled.

I made a mistake for the allValuesFrom restriction, here's the corrected version:

select $s $p $x subquery (
select $instance $t $x subquery (
select
from <...#testexclude>
where exclude($instance $t $x))
from <...#testexclude>
where $s $p $instance and $t <tucana:#is> <rdf:type>
and exclude($instance $t $x))
from <...#testexclude>
where $r <rdf:type> <owl:Restriction>
and $r <owl:onProperty> $p
and $r <owl:allValuesFrom> $x
and $s $p $o2
order by $s ;


I also recently read a paper, "OWL Lite- Reasoning with Rules", which clears up the difference between range and allValuesFrom: "[it] takes into account not only the property but also the domain of the statement. For example, you can say that Stefan manages Wolf, and Wolf is of type Student entails that Stefan is of class Advisor." That's also missing from the query but that's an intentional omission. :-)

Monday, August 30, 2004

Using Kowari's QueryHandler

I've recently been putting builds of Kowari 1.0.5 onto SF's CVS server. This has lead to some good feedback; after some initial bug fixing. It's focused what to work on and the practical up shot is that you can now do RDQL queries (found in org.kowari.store.jena):

RdqlQuery q = new RdqlQuery("select ?x ?y ?z WHERE (?x ?y ?z)");
q.setSource(model);
QueryExecution qe = new KowariQueryEngine(q);
QueryResults results = qe.exec();

This converts the Jena-based query object to Kowari objects, which are then processed by Kowari's query layer and returned as Jena objects.

The only query that it currently doesn't handle, although it soon will, is repeated variables in the constraint like:

select ?x
where (?x ?x ?x)

There's also recent work on client/server JRDF and Jena, N3 input and output, file upload/download from an iTQL client to/from the server, NOT (unstated) constraint, and other improvements.