Paul recently started talking about the history of trans and walk in Kowari/Mulgara. I did a bit of searching in old Tucana emails and feature requests.
I also vaguely remember trying to create an abstract class for walk and the base transitive classes (like the Exhaustive and Direct transitive closure classes). And one of the feature requests also mentions the idea of "set depth" to reduce the recursion on unbounded traversal which would also be handy for the shortest path between two nodes.
Showing posts with label paul gearon. Show all posts
Showing posts with label paul gearon. Show all posts
Tuesday, June 26, 2007
Tuesday, August 22, 2006
Semantic Web News
- Paul announced Mulgara 1.0 which has JRDF 0.3.4.3 (hopefully I'll get a chance to upgrade it to 0.4 now it's Java 1.5). Danny picked up on this and JRDF SPARQL GUI releases.
- TopBraid Composer integrates D2RQ. Featuring Google maps integration and RDF Graph visualization. The documentation has more information on it's use of Jena and Pellet and allows ontology development using the Eclipse framework.
- making sesame2 SPARQL protocol conformant Some thoughts on SPARQL support in Sesame and comments on other SPARQL implementations.
- Not really related to the Semantic Web but via the AI3 blog, Data Visualization at Warp Speed, a brilliant presentation on improvements in health and economic development through the last century through visualization and frantic explanation.
Tuesday, January 10, 2006
Northrop Grumman Killing Kowari?
Resignation from Kowari due to Northrop Grumman Letter "It is with sincere regret that I resigned as an administrator and developer of the Kowari Metastore."
"Northrop Grumman's position seems to be that they "purchased all rights associated with the Kowari software", a position not reconcilable with their continued release of the software under the Mozilla Public License, version 1.1."
So it would seem that Northrop, for whatever reason, dislikes Kowari's existence.
To get David and then Tucana (whose IP was eventually sold to Northrop) to start and continue with an open source version of TKS (the closed source version of the RDF database) was my little hobby horse (I fought fairly long and hard you could say). The idea was to get the Semantic Web bootstrapped and to create a value network around Semantic Web technology (I got this originally from "A critical look at object-orientation"). It is deeply disappointing for me to see it (the value network and Kowari) threatened in this way.
Like any OS project it is the ecosystem that is built around Kowari that makes it work. With so many contributions, just how much of the total code does Northrop actually own the copyright to? By accepting changes it isn't following the dual licencing model used by the owners of Berkley DB. To me Kowari was always using open source as a development and distribution model. Already, we have some contributors threatening to pull out their code (more at, "Disturbing News From Kowari").
Having administrators leave due to Northrop's actions is a good way to kill the project. Hopefully, Northrop's bullying will get a lot of interest - it's good to see, for instance, that this has been picked by aggregators like Topix's Aerospace and Defense.
The other posting David W linked to: Is Northrup Grumman Smushing Kowari?, "But I’d like to know what the status of Kowari is going to be, open source or not, because there are clients for whom Kowari might be the right choice, and we won’t bet on a dog that Northrup Grumman seems determined to publicly beat to death.
Relying on Kowari is now not prudent, given our obligation to do our best for our clients; but it’s also bad for Semantic Web uptake in the US federal government, and that’s something Northrup Grumman should think very carefully about."
Update: Paul and Andrae have also added their thoughts.
Update 2: David has published the letter from Northrop's lawyers. Danny posts, "...clearly if the current Kowari is MPL’d they can’t stop other developers working on the system."
"Northrop Grumman's position seems to be that they "purchased all rights associated with the Kowari software", a position not reconcilable with their continued release of the software under the Mozilla Public License, version 1.1."
So it would seem that Northrop, for whatever reason, dislikes Kowari's existence.
To get David and then Tucana (whose IP was eventually sold to Northrop) to start and continue with an open source version of TKS (the closed source version of the RDF database) was my little hobby horse (I fought fairly long and hard you could say). The idea was to get the Semantic Web bootstrapped and to create a value network around Semantic Web technology (I got this originally from "A critical look at object-orientation"). It is deeply disappointing for me to see it (the value network and Kowari) threatened in this way.
Like any OS project it is the ecosystem that is built around Kowari that makes it work. With so many contributions, just how much of the total code does Northrop actually own the copyright to? By accepting changes it isn't following the dual licencing model used by the owners of Berkley DB. To me Kowari was always using open source as a development and distribution model. Already, we have some contributors threatening to pull out their code (more at, "Disturbing News From Kowari").
Having administrators leave due to Northrop's actions is a good way to kill the project. Hopefully, Northrop's bullying will get a lot of interest - it's good to see, for instance, that this has been picked by aggregators like Topix's Aerospace and Defense.
The other posting David W linked to: Is Northrup Grumman Smushing Kowari?, "But I’d like to know what the status of Kowari is going to be, open source or not, because there are clients for whom Kowari might be the right choice, and we won’t bet on a dog that Northrup Grumman seems determined to publicly beat to death.
Relying on Kowari is now not prudent, given our obligation to do our best for our clients; but it’s also bad for Semantic Web uptake in the US federal government, and that’s something Northrup Grumman should think very carefully about."
Update: Paul and Andrae have also added their thoughts.
Update 2: David has published the letter from Northrop's lawyers. Danny posts, "...clearly if the current Kowari is MPL’d they can’t stop other developers working on the system."
Saturday, December 31, 2005
RETE Rebuked
A recent blog I started reading after criticizing SPARQL. This time its criticizing an entry Paul made about the scalability of the RETE algorithm: "I am a bit confused by the statement that RETE does not scale. This is contrary to a mountain of papers by researchers and developers around the world. From the paragraph, a couple of things come to mind. "(loading indexes) does not need to be done often" tells me the author doesn't understand the purpose and goal of RETE. RETE was designed to solve machine learning problems where data changes rapidly and reasoning is a continuous process. What the author wants is something closer to BitMap indexes used in OLAP products."
A lot of the points raised, like RETE being for changing data, are mentioned in a previous post under Meeting. Drools was chosen as a starting to point to see what kind of system needed to be developed in Kowari.
Also mentioned, bitmap indexing: "Given Tucana is indexing everything, they might as well adapt Bitmap indexing and get better than linear performance. The problem described by the blog is a well understood problem in the OLAP world."
An previous entry, "Relational theory, RETE and Derby" points to some interesting articles about bitmap indexes (available in Oracle 9) and high scalability requirements: "In a large financial institution like a mutual fund company, they may have 1-20 million customers. If each customer has an average of 20-30 positions (aka specific holding of an equity) that means the potential dataset for firm wide compliance rule could involve 20million+ rows. Doing this within 2-5 seconds is rather hard, so it requires using lots of different techniques. In the extreme cases, a company might have 20 million accounts, which means the potential dataset is 600 million rows."
From the OTN article: "B-tree indexes are usually used when columns are unique or near unique; bitmap indexes should be used, or at least considered, in all other cases. While you would not generally use a b-tree index when retrieving 40 percent of the rows of a table, a bitmap index is often still faster than doing a full table scan. This is seemingly in violation of the 80/20 rule, which is to generally use an index when retrieving 20 percent or less of the rows and do a full table scan when retrieving more. Bitmap indexes are smaller and work differently from the 80/20 rule. You can effectively use bitmap indexes even when retrieving large percentages (20 to 80 percent) of a table. Bitmaps can also be used to retrieve conditions based on nulls (since nulls are also indexed) and for "not equal" conditions."
It would appear that this would be suitable for predicate indexation but not generally as both subjects and objects are near unique.
A lot of the points raised, like RETE being for changing data, are mentioned in a previous post under Meeting. Drools was chosen as a starting to point to see what kind of system needed to be developed in Kowari.
Also mentioned, bitmap indexing: "Given Tucana is indexing everything, they might as well adapt Bitmap indexing and get better than linear performance. The problem described by the blog is a well understood problem in the OLAP world."
An previous entry, "Relational theory, RETE and Derby" points to some interesting articles about bitmap indexes (available in Oracle 9) and high scalability requirements: "In a large financial institution like a mutual fund company, they may have 1-20 million customers. If each customer has an average of 20-30 positions (aka specific holding of an equity) that means the potential dataset for firm wide compliance rule could involve 20million+ rows. Doing this within 2-5 seconds is rather hard, so it requires using lots of different techniques. In the extreme cases, a company might have 20 million accounts, which means the potential dataset is 600 million rows."
From the OTN article: "B-tree indexes are usually used when columns are unique or near unique; bitmap indexes should be used, or at least considered, in all other cases. While you would not generally use a b-tree index when retrieving 40 percent of the rows of a table, a bitmap index is often still faster than doing a full table scan. This is seemingly in violation of the 80/20 rule, which is to generally use an index when retrieving 20 percent or less of the rows and do a full table scan when retrieving more. Bitmap indexes are smaller and work differently from the 80/20 rule. You can effectively use bitmap indexes even when retrieving large percentages (20 to 80 percent) of a table. Bitmaps can also be used to retrieve conditions based on nulls (since nulls are also indexed) and for "not equal" conditions."
It would appear that this would be suitable for predicate indexation but not generally as both subjects and objects are near unique.
Monday, September 12, 2005
Indices and Refactoring JRDF
Staring back at me today is something that Paul probably saw last year. But it's something that I found pretty cool because initially it looked like a mistake. Almost as cool as Tom starting to get SPARQL going.
Basically, there are three indices in JRDF: (0,1,2), (1,2,0) and (2,0,1). With the 0 equating to the subject, the 1 the predicate and 2 to the object.
Looking at a refactoring between these internal indices and triples (ordered by subject, predicate and object) there are these three mappings:
* (0,1,2) maps to (s,p,o) using (0,1,2),
* (1,2,0) maps to (s,p,o) using (2,0,1), and
* (2,0,1) maps to (s,p,o) using (1,2,0).
Using the second index as an example, it means that the internal representation maps the third element to the subject, the first element to predicate and the second element to the object.
The surprising thing I noticed, beside the other two indices not being tested correctly, was that the other indices ((2,1,0), (1,0,2), (0,2,1) - these are mentioned in Paul's post) don't seem to have this property. They all seem to map (s,p,o) like (0,1,2) - i.e. using themselves. Whereas, (1,2,0) mapped to itself gives (2,0,1) and (2,0,1) mapped to itself gives (1,2,0).
Basically, there are three indices in JRDF: (0,1,2), (1,2,0) and (2,0,1). With the 0 equating to the subject, the 1 the predicate and 2 to the object.
Looking at a refactoring between these internal indices and triples (ordered by subject, predicate and object) there are these three mappings:
* (0,1,2) maps to (s,p,o) using (0,1,2),
* (1,2,0) maps to (s,p,o) using (2,0,1), and
* (2,0,1) maps to (s,p,o) using (1,2,0).
Using the second index as an example, it means that the internal representation maps the third element to the subject, the first element to predicate and the second element to the object.
The surprising thing I noticed, beside the other two indices not being tested correctly, was that the other indices ((2,1,0), (1,0,2), (0,2,1) - these are mentioned in Paul's post) don't seem to have this property. They all seem to map (s,p,o) like (0,1,2) - i.e. using themselves. Whereas, (1,2,0) mapped to itself gives (2,0,1) and (2,0,1) mapped to itself gives (1,2,0).
Wednesday, August 31, 2005
New Kowari Soon
Kowari v1.1 Pre-release Feature Peek "The v1.1 release of Kowari will include the following features:
o The addition of Paul Gearon's Krule (pronounced "cruel") rules engine. Paul has fast and scalable RDFS support implementated (with the exception of XSD data typing) and is working on implementing more rules for an OWL subset.
o Rewriting of ConstraintExpressions and ModelExpressions to collapse Resolver-specific constraints that should be evaluated together into compound constraints (query rewriting) and permitting Resolvers to advise query engine of the evaluation order of constraints to ensure binding of required variables (query annotation). These features were added by Andrae Muys and Simon Raboczi and will form the basis for a cleaner query language syntax and cleaner integration of external Resolvers, such as GIS systems.
o The removal of Jena support. Kowari v1.0 implemented a Jena storage backend and included an RDQL query language API. This has caused difficulties in maintenance and support as the two projects have gone in different directions. Kowari will remove (not deprecate) support for Jena as of the v1.1 release.
o Support for compiling Kowari under Java 1.5. The Kowari team will be moving to Java version 1.5 as the canonical build version, replacing Java 1.4.2."
Via, Kowari v1.1 Pre-release Feature Peek.
Also, new MindSwap Kowari Library part of PhotoStuff.
o The addition of Paul Gearon's Krule (pronounced "cruel") rules engine. Paul has fast and scalable RDFS support implementated (with the exception of XSD data typing) and is working on implementing more rules for an OWL subset.
o Rewriting of ConstraintExpressions and ModelExpressions to collapse Resolver-specific constraints that should be evaluated together into compound constraints (query rewriting) and permitting Resolvers to advise query engine of the evaluation order of constraints to ensure binding of required variables (query annotation). These features were added by Andrae Muys and Simon Raboczi and will form the basis for a cleaner query language syntax and cleaner integration of external Resolvers, such as GIS systems.
o The removal of Jena support. Kowari v1.0 implemented a Jena storage backend and included an RDQL query language API. This has caused difficulties in maintenance and support as the two projects have gone in different directions. Kowari will remove (not deprecate) support for Jena as of the v1.1 release.
o Support for compiling Kowari under Java 1.5. The Kowari team will be moving to Java version 1.5 as the canonical build version, replacing Java 1.4.2."
Via, Kowari v1.1 Pre-release Feature Peek.
Also, new MindSwap Kowari Library part of PhotoStuff.
Tuesday, June 28, 2005
Northrop Buys Tucana and Continues Kowari
Northrop Grumman Buys the Tucana Knowledge Server "Stunningly for a company their size, Northrop has not only agreed to support Kowari but rushed to do so. I certainly didn't expect a US federal systems integrator to "get" Open Source Software, but times have clearly changed. Their senior managers have made a legitimate effort to figure out the licensing and how to make it work within their business model. I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
A Grumman employee? Or a Thoughtworker with inside knowledge? Building an Agile Enterprise "I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
Update: Paul's doing the handover
A Grumman employee? Or a Thoughtworker with inside knowledge? Building an Agile Enterprise "I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
Update: Paul's doing the handover
Friday, November 19, 2004
More Working Notes
Andrae is out doing both Paul and I with two blogs: Etymon which has links to papers, interesting articles and the like and Circumlocute which is similar to Working notes.
Subscribe to:
Posts (Atom)