toward SPARQL CR, PR, REC, press releases... schedule review "...the valueTesting and sort issues continue to act like difficult-to-drain swamps...The objection from Network Inference saying that SPARQL should have used XQuery remains...I wanted to have a much smaller spec (without SOURCE/GRAPH etc.) and finish earlier... since we didn't meet the April CR milestone, I'm pretty much convinced that what I want and what this WG wants are pretty different..."
RDF DAWG Issues List: sort, "...re-opened 2005-08-03 following comment ORDER with IRIs." (also, extending < operator and lexical vs value space), valueTesting, disjunction, countAggregate , "these are complicated in RDF due to open world notions of equality and inequality.", and bnodeRef, "Users seem to find it useful to refer to bnodes given by the server, though the scope of a bnode is usually one lexical graph."
A look at implementing SPARQL using Lisp: "...parts of the grammar weren't designed to be easily parsed; it has a number of features that are clearly user-focused, such as optional dots. (If it were designed to be easily parsed, it would be s-expressions from the start...)...I have seen objections made to UNION, on the grounds that implementation is difficult. I'd like to nip that one in the bud; I wrote twinql in a few weeks, not full-time, inventing algorithms as I went, and UNION wasn't difficult." Alternatively, I guess it could've been XML.
A response to "evaluating SPARQL w.r.t an RDF query language survey" which lists how you can (or cannot) perform queries that: "Return the labels of all topics that are not titles of publications", "Count the number of authors of a publication" and "Return all publications where the page number is the integer value 8."
Monday, October 10, 2005
Just for Fun
-Ofun "One of the key realizations of modern Internet projects (the oft-quoted Web 2.0) is that on the whole, your users can be trusted. The key is that the users also need to have the tools needed to repair any damage the tiny minority may cause. For a development project, modern version control systems can give you "anarchy with an audit trail". If something does go wrong (intentionally or more likely accidentally), it's easy for any other developer to identify and fix or revert the problem. Having this safety net allows the project to run full-bore without time-wasting process getting in the way, and without undue worry that code quality will suffer...before atomic changesets and quality merge tools, it was extremely difficult to roll back a single change made at some point in the past; now it is much easier to do so. And without a proper test suite, it's hard to tell if a change broke the code in the first place...skill ladders are part of the very definition of fun. True passion and community-building rarely develop around a project that doesn't have such a ladder."
Via Slashdot.
Via Slashdot.
Sunday, October 09, 2005
JRDF 0.3.4.1 Out
Now available for download.
Some thoughts related to this:
* To ensure valid refactorings and bug fixes it seems sensible to go back and do some things properly. It means that the next highest priority will have to be writing an NTriples parser to validate the RDF Test cases. The SableCC grammar for NTriples is already done.
* While this is a good start, it's not nearly fine grained enough to actually write the code and have fast running unit tests. Tom has done some of this test driving a parser with his SPARQL work although I suspect I may do it a little differently (using EasyMock).
* Graph equality and isomorphism on graphs is a pain - which is to say that it's complicated with blank nodes (see the test cases). In "Matching RDF Graphs it lists possible ways to do mappings, including a blank node may map to a labelled node. I know that in Kowari loading the sample FOAF files twice into the same graph adds the blank nodes twice. While this seems like a mistake it is the fastest way to load a graph. Maybe implementing different loading modes (no duplicate checking, blank node to blank node and blank node to labelled node) and signing grahs are two ways to help this.
Some thoughts related to this:
* To ensure valid refactorings and bug fixes it seems sensible to go back and do some things properly. It means that the next highest priority will have to be writing an NTriples parser to validate the RDF Test cases. The SableCC grammar for NTriples is already done.
* While this is a good start, it's not nearly fine grained enough to actually write the code and have fast running unit tests. Tom has done some of this test driving a parser with his SPARQL work although I suspect I may do it a little differently (using EasyMock).
* Graph equality and isomorphism on graphs is a pain - which is to say that it's complicated with blank nodes (see the test cases). In "Matching RDF Graphs it lists possible ways to do mappings, including a blank node may map to a labelled node. I know that in Kowari loading the sample FOAF files twice into the same graph adds the blank nodes twice. While this seems like a mistake it is the fastest way to load a graph. Maybe implementing different loading modes (no duplicate checking, blank node to blank node and blank node to labelled node) and signing grahs are two ways to help this.
Friday, October 07, 2005
JRDF 0.3.4.1 Soon
So the JRDF 0.3.4 was replaced last night due to a problem in the Ant build script creating it with source instead of classes - so much for QA.
Following that little discovery a few other problems related to RDF/XML parsing have been fixed: Does not resolve relative URIs correctly and Stacked xml:base directives with relative URIs not processed. The first is due to the difference between the way Java's URI classs resolves and the way its defined in RDF/XML. In section 5.2 of RFC2396 says: "...any characters after the last (right-most) slash character, if any, are excluded." and the RDF/XML specification says: "These specifications do not specify an algorithm for resolving a fragment identifier alone...The empty string is transformed into an RDF URI reference by substituting the in-scope base URI."
The second is a fix from Sesame's RIO.
Following that little discovery a few other problems related to RDF/XML parsing have been fixed: Does not resolve relative URIs correctly and Stacked xml:base directives with relative URIs not processed. The first is due to the difference between the way Java's URI classs resolves and the way its defined in RDF/XML. In section 5.2 of RFC2396 says: "...any characters after the last (right-most) slash character, if any, are excluded." and the RDF/XML specification says: "These specifications do not specify an algorithm for resolving a fragment identifier alone...The empty string is transformed into an RDF URI reference by substituting the in-scope base URI."
The second is a fix from Sesame's RIO.
Wednesday, October 05, 2005
Hypertime
Writeboard "Every time you save an edit a new version is created and linked in the sidebar. This allows you to write without fear of deleting something, overwriting something, or losing a better version of the document from last week."
Seems very similar to "Reasoning behind the OSMIC proposal: MODELS OF TIME, BACKTRACK AND GROUPWARE"
Seems very similar to "Reasoning behind the OSMIC proposal: MODELS OF TIME, BACKTRACK AND GROUPWARE"
Out meta-ering the competition
Ning - a meta social app platform ""So after being offered a large number of social media applications, we are now into the meta-framework to build social media applications. The notion is interesting: as we have come to expect that any consumer application will include some element of social networking, collaborative filtering, tagging, etc., Ning has the first shot at claiming platform status in the social phenomena by offering building consistent building blocks (though wikis could probably claim anteriority)."...What it does do is make APIs an even more important feature for social apps to have, because Ning users will potentially have less need to visit the websites of Flickr, del.icio.us, 43Things and others. Ning will drive customers to those other services, but via the APIs. "
I always seem to quote people who quote other people, seems apt in this case.
I always seem to quote people who quote other people, seems apt in this case.
Tuesday, October 04, 2005
Gun
Google and Sun together "Google and Sun Microsystems will hold a press conference on Tuesday at which they're expected to announce a collaboration to bring StarOffice productivity applications to Google users..."Imagine StarOffice running on the desktop, and Google perfecting the [file synchronization]," said Edwards. "Then you have your collaboration space carved out immediately for you, and Google is hosting it.""
Does it even make sense for a company that is selling open standard, network centric, highly scalable applications on Unix combines with a company that invented Java?
Imagine a host of GApps (GMail, GCal, GIM, GWorld, etc.) working on your desktop.
Via Update: Sun Marries Google? and a smidge more Sun turns up volume against Windows Vista and Office 12 (links to The Value in Volume).
Does it even make sense for a company that is selling open standard, network centric, highly scalable applications on Unix combines with a company that invented Java?
Imagine a host of GApps (GMail, GCal, GIM, GWorld, etc.) working on your desktop.
Via Update: Sun Marries Google? and a smidge more Sun turns up volume against Windows Vista and Office 12 (links to The Value in Volume).
Monday, October 03, 2005
The Future is Google Ads Everywhere (powered by SPARQL)
Scoble ? RDF "SPARQL is an answer to the question “What if I want to do SQL-like querying when I know perfectly well that everybody will be using their own incompatible database schema?” I’ve been a SemWeb skeptic, but I look at SPARQL and I think: Suppose you could assemble a ton of property-value pairs about web sites, and suppose on the front end you could build a nice responsive query page that allowed you to compose queries like Scoble’s hotel search; well then, SPARQL would be more or less exactly what you need to bridge the gap. Hey, isn’t Guha’s Alpiri project more or less that back-end? And isn’t Guha working at Google now?"
And what would we do with this perfect engine and ubiquitous bandwidth - ads. Author: Google's Patents Reveal Strategy To Beat Microsof "In Arnold’s analysis, he said some filings in the patent portfolio point to an accelerated use of high-speed fiber and wireless that could be used to deliver Google technology...An ultimate goal of the firm is to deliver completely individualized ads to users."
And what would we do with this perfect engine and ubiquitous bandwidth - ads. Author: Google's Patents Reveal Strategy To Beat Microsof "In Arnold’s analysis, he said some filings in the patent portfolio point to an accelerated use of high-speed fiber and wireless that could be used to deliver Google technology...An ultimate goal of the firm is to deliver completely individualized ads to users."
Friday, September 30, 2005
Evolution not creation
Another speech by Adam Bosworth He talked mostly on a TDD/Agile theme where applications are developed in 2 week iterations based directly on customer usage. This offers a better alternative to software architects spending years in seclusion developing the be-all and end-all API and developers hoping that they will get what they need. He cited Windows as an example of this approach. Unsurpisingly, this leads to more tightly coupled code as reported in the recent article about Microsoft rewriting Windows by the Wall Street Journal.
Bosworth calls this approach intelligent reaction rather than intelligent design. Google, eBay and Salesforce were given as examples of this application driven rather than API driven development.
Relating this back to RDF and agile databases, the experience at Salesforce is that the most significant amount of effort is spent customising the data not in other areas such as the user-interface or processing logic. While he doesn't give a reason for data over processing he does say that customised user interfaces are too expensive because they require more user training.
A transcript of some of the talk is available at: Bosworth: The new model is, 'run like mad'.
Bosworth calls this approach intelligent reaction rather than intelligent design. Google, eBay and Salesforce were given as examples of this application driven rather than API driven development.
Relating this back to RDF and agile databases, the experience at Salesforce is that the most significant amount of effort is spent customising the data not in other areas such as the user-interface or processing logic. While he doesn't give a reason for data over processing he does say that customised user interfaces are too expensive because they require more user training.
A transcript of some of the talk is available at: Bosworth: The new model is, 'run like mad'.
JVM Dynamic Language Support
invokedynamic: New Java Bytecode for the Dynamics "The new byte code, invokedynamic , is coming to a JSR near you very soon....the verifier won’t insist that the type of the target of the method invocation (the receiver, in Smalltalk speak) be known to support the method being invoked, or that the types of the arguments be known to match the signature of that method. Instead, these checks will be done dynamically."
From a previous article, Pluggable Types "...Ruby (that's basically Smalltalk with a Perl style syntax)..."
From a previous article, Pluggable Types "...Ruby (that's basically Smalltalk with a Perl style syntax)..."
Saturday, September 24, 2005
Supping from the Information Firehose
* Oracle 10g Support for RDF "There are procedures provided to search the RDF models with a language that looks like a bit like SPARQL...The search is pretty nice because you can join your search on standard SQL tables, thus combining both your triple model and relational models together.
Oracle also has Rule support, in the form of Rule Indexes. A builtin set of rules for RDF Schema (RDFS) semantics is provided, which is a very nice touch. You are also free to create your own rules, both with the query pattern matching and filters."
* An Atom Store With links to people wanting to create a non-SQL store for Atom, including the very good Bosworth's Web of Data (original here) which is largely about why new databases should not be like Oracle but should distribute the processing (no views, triggers, etc).
* RDFAuthor does SPARQL "Damian is working on a new version of RDFAuthor that generates SPARQL queries (instead of the older Squish notation). It can also (not sure which protocol(s)) get results from a query service." I liked the original RDFAuthor - good to see an update is underway.
* Wrapping rdflib's Graph around a 4RDF Model "I wrote a 4Suite RDF model backend for rdflib, that allows the wrapping of Graph around a live 4Suite RDF model. Finally, I used this backend to execute a sparql-p query..."
* Secrets of lightweight development success, Part 7: Java alternatives Closures, Continuations, Metaprogramming and Reflection: "In short, the Java language just isn't a very productive applications language. The founders made some wise compromises to wrestle control away from C++, but we're starting to pay for those compromises."
* ONJava 2005 Reader Survey Results, Part 1 "Eclipse (76 percent), NetBeans (21 percent), None (17 percent), IntelliJ (13 percent)...JBoss (38 percent), None (28 percent), WebSphere (21 percent), WebLogic (20 percent)"
* Language Innovation: C# 3.0 explained "...most of the features of C# 3.0 are, arguably, nothing but syntactic sugar designed to make programming more productive..." One of the things that was (originally) good about Java was the lack of syntax (I thought anyway!).
* KVM over IP
Oracle also has Rule support, in the form of Rule Indexes. A builtin set of rules for RDF Schema (RDFS) semantics is provided, which is a very nice touch. You are also free to create your own rules, both with the query pattern matching and filters."
* An Atom Store With links to people wanting to create a non-SQL store for Atom, including the very good Bosworth's Web of Data (original here) which is largely about why new databases should not be like Oracle but should distribute the processing (no views, triggers, etc).
* RDFAuthor does SPARQL "Damian is working on a new version of RDFAuthor that generates SPARQL queries (instead of the older Squish notation). It can also (not sure which protocol(s)) get results from a query service." I liked the original RDFAuthor - good to see an update is underway.
* Wrapping rdflib's Graph around a 4RDF Model "I wrote a 4Suite RDF model backend for rdflib, that allows the wrapping of Graph around a live 4Suite RDF model. Finally, I used this backend to execute a sparql-p query..."
* Secrets of lightweight development success, Part 7: Java alternatives Closures, Continuations, Metaprogramming and Reflection: "In short, the Java language just isn't a very productive applications language. The founders made some wise compromises to wrestle control away from C++, but we're starting to pay for those compromises."
* ONJava 2005 Reader Survey Results, Part 1 "Eclipse (76 percent), NetBeans (21 percent), None (17 percent), IntelliJ (13 percent)...JBoss (38 percent), None (28 percent), WebSphere (21 percent), WebLogic (20 percent)"
* Language Innovation: C# 3.0 explained "...most of the features of C# 3.0 are, arguably, nothing but syntactic sugar designed to make programming more productive..." One of the things that was (originally) good about Java was the lack of syntax (I thought anyway!).
* KVM over IP
Labels:
adam bosworth,
c#,
closures,
continuations,
ide,
j2ee,
java,
kvm,
mysql,
oracle,
programming languages,
rdf,
sparql
Wednesday, September 21, 2005
What do you do all day?
Code Is Not An Asset "Since software code is not an asset, but rather a liability, the more we can reduce the deadwood, the better off we are...It is definitely possible to deliver high level of functionality, interactivity and sophistication by utilizing only a portion of code that would normally be used if we stick to the old school (morecode, or more LOC). And that’s a desirable thing."
Monday, September 19, 2005
What's a Unit Test?
A Set of Unit Testing Rules "A test is not a unit test if:
* It talks to the database
* It communicates across the network
* It touches the file system
* It can't run at the same time as any of your other unit tests
* You have to do special things to your environment (such as editing config files) to run it."
"If you write code in a way which separates your logic from OS and vendor services, you not only get faster unit tests, you get a ‘binary chop’ that allows you to discover whether the problem is in your logic or in the things are you interfacing with."
* It talks to the database
* It communicates across the network
* It touches the file system
* It can't run at the same time as any of your other unit tests
* You have to do special things to your environment (such as editing config files) to run it."
"If you write code in a way which separates your logic from OS and vendor services, you not only get faster unit tests, you get a ‘binary chop’ that allows you to discover whether the problem is in your logic or in the things are you interfacing with."
Linkaholic
* Questioning RDF "'Now I hear the argument that one does not need to know hedge automata to use RELAX NG, and all that, but I don't think it applies in the case of RDF. In RDF, the model semantics are the primary reason for coming to the party. I don't see it as an optional formalization. Maybe I'm wrong about that and it's the need to write a query language for RDF (hardly typical for the Web punter) that is causing me to gurgle in the muck.'...I definitely think there is some merit to disconnecting RDF from the Semantic Web and seeing if it can hang on its own from that perspective...I've wondered if there is similar usefulness lurking within RDF once it loses its Semantic Web baggage."
* An early look at JUnit 4 If you don't like TestNG or JUnit 4 what do you do? I like the "failXXX() throws FooBarException" better. The annotations just don't do it for me.
* XML Virtual Machines "You can target XUL for deployment to the Mozilla platform, today. XULRunner now means you can develop and test away in double quick time . Yes, Mozilla is a platform (a XUL visual forms builder could be a game changer on the client, as well saving some folks I know a lot of typing ;). And if you need a thicker-than-that client, then consider the OpenOffice platform."
* The Future of Mobility is Linux "The latest news I’ve seen about the Nokia 770 is that it’s going to have a host of applications ready for it at launch, including VoIP software, streaming media, chat applications, Doom, etc. The thing that’s so amazing about this is that the 770 is essentially the *same exact hardware* that’s on my Nokia 6680, yet the development pace for the 770 is way more rapid. In addition, there’s at least a half a dozen blogs and bloggers dedicated to the device, and it hasn’t even launched yet. This shows the power of an open environment and the draw of Linux and its fans."
* An early look at JUnit 4 If you don't like TestNG or JUnit 4 what do you do? I like the "failXXX() throws FooBarException" better. The annotations just don't do it for me.
* XML Virtual Machines "You can target XUL for deployment to the Mozilla platform, today. XULRunner now means you can develop and test away in double quick time . Yes, Mozilla is a platform (a XUL visual forms builder could be a game changer on the client, as well saving some folks I know a lot of typing ;). And if you need a thicker-than-that client, then consider the OpenOffice platform."
* The Future of Mobility is Linux "The latest news I’ve seen about the Nokia 770 is that it’s going to have a host of applications ready for it at launch, including VoIP software, streaming media, chat applications, Doom, etc. The thing that’s so amazing about this is that the 770 is essentially the *same exact hardware* that’s on my Nokia 6680, yet the development pace for the 770 is way more rapid. In addition, there’s at least a half a dozen blogs and bloggers dedicated to the device, and it hasn’t even launched yet. This shows the power of an open environment and the draw of Linux and its fans."
Adding more Layers
Crisis "What if we jilted the ugly sisters of rdf:Bag, rdf:Alt and rdf:Alt and took reification out back and shot it? How many tears would be shed?
What if we junked classes, domains and ranges? Would anyone notice? The key concept in RDF is the relationship, the property.
The result would be a subset of RDF, RDF-lite perhaps. All instances of RDF-lite would be valid RDF-full but the converse couldn’t be true. Sparql would still work and so, I suspect, would the OWL machinery despite the omission of classes. RDF diffs would be trivial without blank nodes allowing efficient synchronisation of triple stores. Signing of triples would also be possible without requiring the hoops of canonicalisation to be jumped through."
Sunday afternoon blather "So (ramble nearly over) basically I don’t think there’s any need to define a subset or simplification of RDF because it’s already layered. You don’t need reification, ontological inferences don’t use them. If all you want to do is publish HTML with a bit of explicit data embedded, or write an aggregator that understands “rel” then fine. If you’re doing good Web stuff, you’re still helping the Semantic Web. What SPARQL brings is a low barrier to entry into the ideas, and a low barrier to development of true Web applications. I do believe SPARQL has the potential to be explosive because it makes the Semantic Web that much more agile."
This got me thinking about a conversation with Tom about the Graph interface in JRDF. It allows you to get TripleFactory (for Collections, reification, etc.) and GraphElementFactory for creating graph elements (nodes and err triples - which is now moved to TripleFactory in 0.4). It makes sense to make these decoupled to provide a tighter interface (for understandability) and less to mock out or stub out when testing.
What if we junked classes, domains and ranges? Would anyone notice? The key concept in RDF is the relationship, the property.
The result would be a subset of RDF, RDF-lite perhaps. All instances of RDF-lite would be valid RDF-full but the converse couldn’t be true. Sparql would still work and so, I suspect, would the OWL machinery despite the omission of classes. RDF diffs would be trivial without blank nodes allowing efficient synchronisation of triple stores. Signing of triples would also be possible without requiring the hoops of canonicalisation to be jumped through."
Sunday afternoon blather "So (ramble nearly over) basically I don’t think there’s any need to define a subset or simplification of RDF because it’s already layered. You don’t need reification, ontological inferences don’t use them. If all you want to do is publish HTML with a bit of explicit data embedded, or write an aggregator that understands “rel” then fine. If you’re doing good Web stuff, you’re still helping the Semantic Web. What SPARQL brings is a low barrier to entry into the ideas, and a low barrier to development of true Web applications. I do believe SPARQL has the potential to be explosive because it makes the Semantic Web that much more agile."
This got me thinking about a conversation with Tom about the Graph interface in JRDF. It allows you to get TripleFactory (for Collections, reification, etc.) and GraphElementFactory for creating graph elements (nodes and err triples - which is now moved to TripleFactory in 0.4). It makes sense to make these decoupled to provide a tighter interface (for understandability) and less to mock out or stub out when testing.
Lack of Symmetry
Catching up on Journal of Web Semantics: Preprint Server:
* Completeness, decidability and complexity of entailment for RDF Schema and a semantic extension involving the OWL vocabulary "RDFS entailment is decidable, NP-complete, and in P if the target graph does not contain blank nodes...In this paper we extend the notion of RDF graph to allow blank nodes in predicate position and show that the standard entailment rules for RDFS become complete if rule rdfs7 is replaced by a suitably extended rule rdfs7x. In view of the fact that the semantics of RDFS allows blank nodes to refer to properties, ‘generalized RDF graphs’, which allow blank nodes in predicate position, seem to form a more direct abstract syntax for RDFS than RDF graphs."
* OWL-Eu: Adding Customised Datatypes into OWL "1. OWL does not support customised datatypes (except enumerated datatypes)...
2. OWL does not support negated datatypes. For example, ‘all integers but 0’...
3. An OWL DL datatype domain seriously restricts the interpretations of typed literals with unsupported datatype URIrefs..."
"OWL-Eu supports customised datatypes through unary datatype expressions based on unary datatype groups. Intuitively, an unary datatype group extends the OWL datatyping with a hierarchy of supported datatypes."
* Completeness, decidability and complexity of entailment for RDF Schema and a semantic extension involving the OWL vocabulary "RDFS entailment is decidable, NP-complete, and in P if the target graph does not contain blank nodes...In this paper we extend the notion of RDF graph to allow blank nodes in predicate position and show that the standard entailment rules for RDFS become complete if rule rdfs7 is replaced by a suitably extended rule rdfs7x. In view of the fact that the semantics of RDFS allows blank nodes to refer to properties, ‘generalized RDF graphs’, which allow blank nodes in predicate position, seem to form a more direct abstract syntax for RDFS than RDF graphs."
* OWL-Eu: Adding Customised Datatypes into OWL "1. OWL does not support customised datatypes (except enumerated datatypes)...
2. OWL does not support negated datatypes. For example, ‘all integers but 0’...
3. An OWL DL datatype domain seriously restricts the interpretations of typed literals with unsupported datatype URIrefs..."
"OWL-Eu supports customised datatypes through unary datatype expressions based on unary datatype groups. Intuitively, an unary datatype group extends the OWL datatyping with a hierarchy of supported datatypes."
Thursday, September 15, 2005
Really Dynamic Framework
Semantic Rails, Semantic Django: Pushing RDF into MVC "...what Rails-Django would look like if instead of SQL and RDBMS... What if the "M" in MVC were composed of RDF, SPARQL, and a triplestore like Kowari?
Sure, Kowari is probably slower than MySQL, and you probably know SQL a lot better than SPARQL, but RDF is a schemaless data representation thingie. You can start with as little or as much schema as you want or need, and you can use the full expressive power of OWL (which is significantly more powerful than SQL's DDL) when you need it."
Four points are raised in "Semantic MVC".
nodel - HOWTO "Nodel is a kind of application generator. It is a halfway house between the ontomatic, which isn't yet stable enough in implementation or customisable enough in interface, and the nodedb, which was very hardcoded and application specific."
Related to: RDF, the ultimate agile database, Another Agile Database User and Scripting the Semantic Web.
Sure, Kowari is probably slower than MySQL, and you probably know SQL a lot better than SPARQL, but RDF is a schemaless data representation thingie. You can start with as little or as much schema as you want or need, and you can use the full expressive power of OWL (which is significantly more powerful than SQL's DDL) when you need it."
Four points are raised in "Semantic MVC".
nodel - HOWTO "Nodel is a kind of application generator. It is a halfway house between the ontomatic, which isn't yet stable enough in implementation or customisable enough in interface, and the nodedb, which was very hardcoded and application specific."
Related to: RDF, the ultimate agile database, Another Agile Database User and Scripting the Semantic Web.
Bitten by the Software Health Bug
SPARQL support Tom comments about a number of things we have both learned since working at Tucana on making better software: "The unit tests are not unit tests. Most of the tests in Kowari are integration tests, that is, while they may test only a single class, they test that class with the roles of its dependent class filled by real concrete instances, rather than mocks...Andrew and I have been bitten by the agile, quality, TDD and refactoring bug, and having our own small project to play with makes this much easier. It also allows us to enforce higher standards on the codebase as there's only two of us that need to agree at the moment."
"I hope none of this comes off as criticism for Kowari as this wasn't my intention. It's certainly the best triplestore there is, and the project is moving ahead at great guns now. It's basically a reflection that Kowari is too heavyweight to do what we want to do and that both projects have different goals."
The difference in what is a unit test is significant, Martin Fowler's "Mocks Aren't Stubs" is a good introduction - although he remains on the fence. We didn't - we used interaction based testing throughout. We still had the usual integration tests but it also included environment, wiring (Spring), and others. Interaction based testing seems to lead to better software. We even test drove XML, properties, etc.
How the software is better is not in the quality code so much (although I do think the initial number defects is also lower) but in the health of the code. A healthy code base can respond quickly to changes in architecture, requirements, etc. or in other, trendier words - it's more agile. Kent Beck's talk on software health is well summarized here.
Paul O'Keeffe and Greg Davis gave a presentation on Sustainable Software Development - firmly on the interaction testing side. You really do seem to get benefits late in a project with a sustainable, healthy code base.
I think it's covered in, "Using Software Maintainability Models to Track Code Health" but it doesn't seem to be available online. Another paper though, Software Deterioration And Maintainability – A Model Proposal contrasts the difference between "fast hack" and a "controlled update" for example and the need for just the right amount of effort for maintaining software. Finding research on interaction versus state based testing would be interesting.
JRDF still needs refactoring to get into a stage where it is test driven. And test driving isn't always a good tool to develop solutions to problems (to hack, to spike). So there's still a place for that too.
"I hope none of this comes off as criticism for Kowari as this wasn't my intention. It's certainly the best triplestore there is, and the project is moving ahead at great guns now. It's basically a reflection that Kowari is too heavyweight to do what we want to do and that both projects have different goals."
The difference in what is a unit test is significant, Martin Fowler's "Mocks Aren't Stubs" is a good introduction - although he remains on the fence. We didn't - we used interaction based testing throughout. We still had the usual integration tests but it also included environment, wiring (Spring), and others. Interaction based testing seems to lead to better software. We even test drove XML, properties, etc.
How the software is better is not in the quality code so much (although I do think the initial number defects is also lower) but in the health of the code. A healthy code base can respond quickly to changes in architecture, requirements, etc. or in other, trendier words - it's more agile. Kent Beck's talk on software health is well summarized here.
Paul O'Keeffe and Greg Davis gave a presentation on Sustainable Software Development - firmly on the interaction testing side. You really do seem to get benefits late in a project with a sustainable, healthy code base.
I think it's covered in, "Using Software Maintainability Models to Track Code Health" but it doesn't seem to be available online. Another paper though, Software Deterioration And Maintainability – A Model Proposal contrasts the difference between "fast hack" and a "controlled update" for example and the need for just the right amount of effort for maintaining software. Finding research on interaction versus state based testing would be interesting.
JRDF still needs refactoring to get into a stage where it is test driven. And test driving isn't always a good tool to develop solutions to problems (to hack, to spike). So there's still a place for that too.
Monday, September 12, 2005
Indices and Refactoring JRDF
Staring back at me today is something that Paul probably saw last year. But it's something that I found pretty cool because initially it looked like a mistake. Almost as cool as Tom starting to get SPARQL going.
Basically, there are three indices in JRDF: (0,1,2), (1,2,0) and (2,0,1). With the 0 equating to the subject, the 1 the predicate and 2 to the object.
Looking at a refactoring between these internal indices and triples (ordered by subject, predicate and object) there are these three mappings:
* (0,1,2) maps to (s,p,o) using (0,1,2),
* (1,2,0) maps to (s,p,o) using (2,0,1), and
* (2,0,1) maps to (s,p,o) using (1,2,0).
Using the second index as an example, it means that the internal representation maps the third element to the subject, the first element to predicate and the second element to the object.
The surprising thing I noticed, beside the other two indices not being tested correctly, was that the other indices ((2,1,0), (1,0,2), (0,2,1) - these are mentioned in Paul's post) don't seem to have this property. They all seem to map (s,p,o) like (0,1,2) - i.e. using themselves. Whereas, (1,2,0) mapped to itself gives (2,0,1) and (2,0,1) mapped to itself gives (1,2,0).
Basically, there are three indices in JRDF: (0,1,2), (1,2,0) and (2,0,1). With the 0 equating to the subject, the 1 the predicate and 2 to the object.
Looking at a refactoring between these internal indices and triples (ordered by subject, predicate and object) there are these three mappings:
* (0,1,2) maps to (s,p,o) using (0,1,2),
* (1,2,0) maps to (s,p,o) using (2,0,1), and
* (2,0,1) maps to (s,p,o) using (1,2,0).
Using the second index as an example, it means that the internal representation maps the third element to the subject, the first element to predicate and the second element to the object.
The surprising thing I noticed, beside the other two indices not being tested correctly, was that the other indices ((2,1,0), (1,0,2), (0,2,1) - these are mentioned in Paul's post) don't seem to have this property. They all seem to map (s,p,o) like (0,1,2) - i.e. using themselves. Whereas, (1,2,0) mapped to itself gives (2,0,1) and (2,0,1) mapped to itself gives (1,2,0).
Thursday, September 08, 2005
Further Beyond Java
A little bit late but The JavaCast has an interview with Bruce Tate.
Update (15/11): Fixed the link to the MP3.
Update (15/11): Fixed the link to the MP3.
Wednesday, September 07, 2005
Stacks of StAX stacks compared
Streaming APIs for XML Parsers "The performance of the five parsers was measured over a large number of XML documents." BEA, Oracle, SJSXP, XPP3 and Woodstox compared. The last two come out on top, although "...XPP3 is based on XmlPullParser APIs and not JSR-173 compliant. XPP3 is a parsing API that will work with small devices (J2ME compatible)." SJSXP has the advantage that it supports, "...symmetrical bi-directional APIs that can both read and write XML documents using the same representation of XML."
Via StAX parsing performance paper.
Via StAX parsing performance paper.
Tuesday, September 06, 2005
RDF in the USA
America seems to be where it's at for the Semantic Web:
* Work in the USA "...I decided that I should take the job. This is a big deal...To start with, I'm a contractor working remotely with Herzum until I can get a visa to become a full time employee."
* Semantic Web Yahoo! "At the end of September I'll be leaving the Institute for Learning and Research Technology (ILRT) at the University of Bristol in the UK after more than five years of productive and interesting work with RDF and the Semantic Web including developing the Redland RDF libraries. It's been a great place to work at with super people who have been very supportive of the Semantic Web and innovation."
If you ignore those people at Maryland or MIT or whatever there does seem to have been more European interest in the Semantic Web - maybe this is set to change. This is probably what it takes for it to go mainstream.
* Work in the USA "...I decided that I should take the job. This is a big deal...To start with, I'm a contractor working remotely with Herzum until I can get a visa to become a full time employee."
* Semantic Web Yahoo! "At the end of September I'll be leaving the Institute for Learning and Research Technology (ILRT) at the University of Bristol in the UK after more than five years of productive and interesting work with RDF and the Semantic Web including developing the Redland RDF libraries. It's been a great place to work at with super people who have been very supportive of the Semantic Web and innovation."
If you ignore those people at Maryland or MIT or whatever there does seem to have been more European interest in the Semantic Web - maybe this is set to change. This is probably what it takes for it to go mainstream.
Thursday, September 01, 2005
Some Updates
* Ajax Libraries it seems the DWR is best, see demos.
* Web 2.0 "Proponents of the Web 2.0 approach believe that web usage is increasingly oriented toward interaction and rudimentary social networks, which can serve content that exploits network effects with or without creating a visual, interactive web page. In one view, Web 2.0 sites act more as points of presence, or user-dependent web portals, than as traditional websites."
* Clear message for causality "Stenner and co-workers found that although the smooth pulse arrives noticeably earlier through the superluminal medium, the instant at which the "1s" and "0s" begin to differ does not seem to be accelerated."
* Web 2.0 "Proponents of the Web 2.0 approach believe that web usage is increasingly oriented toward interaction and rudimentary social networks, which can serve content that exploits network effects with or without creating a visual, interactive web page. In one view, Web 2.0 sites act more as points of presence, or user-dependent web portals, than as traditional websites."
* Clear message for causality "Stenner and co-workers found that although the smooth pulse arrives noticeably earlier through the superluminal medium, the instant at which the "1s" and "0s" begin to differ does not seem to be accelerated."
Wednesday, August 31, 2005
AJAX and REST
Why I think AJAX breaks REST "The Web is a hypermedia application. According to REST, following a link you should be transferred to a state representation which is associated with a particular URI (in the case of the Web, a URL). As it happens, together with that state representation we may also get some code that our hypermedia document viewer (the browser) executes. Depending on the interactions with the end-user, another program, the environment, etc. the executed code may make use of asynchronous calls to change the represented state that our viewer has loaded. It does so by usually fetching more state, possibly in a REST-friendly way (i.e. HTTP requests). However, during the user-application interactions, the URI remains the same even though the state representation that is rendered on our screens has obviously changed. Now, some AJAX applications may force a change in the URI effectively simulating a state transition but most of the ones I have used do not."
REST is a way to expose services and AJAX is a way to write client software - they can't really effect one another.
AJAX can break the ability to save client state, in exactly the same way that frames do (complete with breaking the back button). AJAX applications may or may not update the URI in the browser to represent the current state it is in, but you can certainly write them so that they do (just like frames).
In a RESTful approach the architecture puts restrictions on the server, not how the client represents its current state:"...each request from client to server must contain all the information necessary to understand the request, and cannot take advantage of any stored context on the server."
Also, AJAX - beyond the buzzwords.
Update: Why I think AJAX breaks REST continued "What I did say is that the use of the AJAX approach to building distributed applications, or even just client-side ones, encourages a move away from the REST principles and that those building distributed applications should be aware of that...We just can't call all the good things out there RESTful."
I take REST principles to mean providing URIs to represent the current state they are in. However, I don't see browsers or client side applications as being RESTful or not. Tabs, frames, pop-ups, running code in Flash, Javascript, and applets; all of these things are not represented by the URI in the toolbar.
The only thing that's guaranteed in REST is that a client changes state when it goes to a URI (the representation of the state) on a server. How it represents that state is entirely up to it.
There are good reasons why everything just isn't represented by a URI, but I do think that more things should be. One is space requirements, imagine GMail storing every state that your mail box is in and providing the ability to go back and forward (very Xanadu perhaps).
REST is a way to expose services and AJAX is a way to write client software - they can't really effect one another.
AJAX can break the ability to save client state, in exactly the same way that frames do (complete with breaking the back button). AJAX applications may or may not update the URI in the browser to represent the current state it is in, but you can certainly write them so that they do (just like frames).
In a RESTful approach the architecture puts restrictions on the server, not how the client represents its current state:"...each request from client to server must contain all the information necessary to understand the request, and cannot take advantage of any stored context on the server."
Also, AJAX - beyond the buzzwords.
Update: Why I think AJAX breaks REST continued "What I did say is that the use of the AJAX approach to building distributed applications, or even just client-side ones, encourages a move away from the REST principles and that those building distributed applications should be aware of that...We just can't call all the good things out there RESTful."
I take REST principles to mean providing URIs to represent the current state they are in. However, I don't see browsers or client side applications as being RESTful or not. Tabs, frames, pop-ups, running code in Flash, Javascript, and applets; all of these things are not represented by the URI in the toolbar.
The only thing that's guaranteed in REST is that a client changes state when it goes to a URI (the representation of the state) on a server. How it represents that state is entirely up to it.
There are good reasons why everything just isn't represented by a URI, but I do think that more things should be. One is space requirements, imagine GMail storing every state that your mail box is in and providing the ability to go back and forward (very Xanadu perhaps).
Never Forget You are the Lowest of the Low
When thinking about French animation or cartoons the first thoughts are Asterix, The Smurfs or Tintin (which is actually Belgian).
Others include:
* Robostory or here.
* Spartakus and the Sun Beneath the Sea or here.
* Le Croc-Note Show.
* Inspector Gadget.
* Ulysses 31.
* Belle and Sebastian.
* Babar the Elephant.
* The Snorks.
* Madeline.
And while not French, Vicky the Viking was a boy.
Others include:
* Robostory or here.
* Spartakus and the Sun Beneath the Sea or here.
* Le Croc-Note Show.
* Inspector Gadget.
* Ulysses 31.
* Belle and Sebastian.
* Babar the Elephant.
* The Snorks.
* Madeline.
And while not French, Vicky the Viking was a boy.
New Kowari Soon
Kowari v1.1 Pre-release Feature Peek "The v1.1 release of Kowari will include the following features:
o The addition of Paul Gearon's Krule (pronounced "cruel") rules engine. Paul has fast and scalable RDFS support implementated (with the exception of XSD data typing) and is working on implementing more rules for an OWL subset.
o Rewriting of ConstraintExpressions and ModelExpressions to collapse Resolver-specific constraints that should be evaluated together into compound constraints (query rewriting) and permitting Resolvers to advise query engine of the evaluation order of constraints to ensure binding of required variables (query annotation). These features were added by Andrae Muys and Simon Raboczi and will form the basis for a cleaner query language syntax and cleaner integration of external Resolvers, such as GIS systems.
o The removal of Jena support. Kowari v1.0 implemented a Jena storage backend and included an RDQL query language API. This has caused difficulties in maintenance and support as the two projects have gone in different directions. Kowari will remove (not deprecate) support for Jena as of the v1.1 release.
o Support for compiling Kowari under Java 1.5. The Kowari team will be moving to Java version 1.5 as the canonical build version, replacing Java 1.4.2."
Via, Kowari v1.1 Pre-release Feature Peek.
Also, new MindSwap Kowari Library part of PhotoStuff.
o The addition of Paul Gearon's Krule (pronounced "cruel") rules engine. Paul has fast and scalable RDFS support implementated (with the exception of XSD data typing) and is working on implementing more rules for an OWL subset.
o Rewriting of ConstraintExpressions and ModelExpressions to collapse Resolver-specific constraints that should be evaluated together into compound constraints (query rewriting) and permitting Resolvers to advise query engine of the evaluation order of constraints to ensure binding of required variables (query annotation). These features were added by Andrae Muys and Simon Raboczi and will form the basis for a cleaner query language syntax and cleaner integration of external Resolvers, such as GIS systems.
o The removal of Jena support. Kowari v1.0 implemented a Jena storage backend and included an RDQL query language API. This has caused difficulties in maintenance and support as the two projects have gone in different directions. Kowari will remove (not deprecate) support for Jena as of the v1.1 release.
o Support for compiling Kowari under Java 1.5. The Kowari team will be moving to Java version 1.5 as the canonical build version, replacing Java 1.4.2."
Via, Kowari v1.1 Pre-release Feature Peek.
Also, new MindSwap Kowari Library part of PhotoStuff.
Light at any Speed
Stoping light in quantum leap "...researchers have slowed the speed of light down from 300,000 kilometres per second to a few hundred metres per second.
“To store the light in there, we turn the second laser beam off. The signal from the first laser beam is trapped inside the crystal. To get the signal out again, we turn the coupling beam on again. We can now store light for seconds, and potentially quite a bit longer,” Dr Sellars said. "
Euro boffins increase speed of light "Luc Thévenaz and his fellow researchers in the Nanophotonics and Metrology laboratory at EPFL said they were able not only to slow light down by a factor of three from its usual speed of 300 million metres per second in a vacuum, but they have also accomplished the considerable feat of speeding it up – effectively making light go faster than the speed of light."
They have "...demonstrated the first all-optical technique to slow light in off-the-shelf optical fibres."
A previous article explaining how superluminal speeds are achieved: "...a pulse of light can have more than one speed because it is made up of light of different wavelengths. The individual waves travel at their own phase velocity, while the pulse itself travels with the group velocity. In a vacuum all the phase velocities and the group velocity are the same. In a dispersive medium, however, they are different because the refractive index is a function of wavelength, which means that the different wavelengths travel at different speeds. Wang and colleagues report evidence for a negative group velocity of -310c, where c (=300 million metres per second) is the speed of light in vacuum."
“To store the light in there, we turn the second laser beam off. The signal from the first laser beam is trapped inside the crystal. To get the signal out again, we turn the coupling beam on again. We can now store light for seconds, and potentially quite a bit longer,” Dr Sellars said. "
Euro boffins increase speed of light "Luc Thévenaz and his fellow researchers in the Nanophotonics and Metrology laboratory at EPFL said they were able not only to slow light down by a factor of three from its usual speed of 300 million metres per second in a vacuum, but they have also accomplished the considerable feat of speeding it up – effectively making light go faster than the speed of light."
They have "...demonstrated the first all-optical technique to slow light in off-the-shelf optical fibres."
A previous article explaining how superluminal speeds are achieved: "...a pulse of light can have more than one speed because it is made up of light of different wavelengths. The individual waves travel at their own phase velocity, while the pulse itself travels with the group velocity. In a vacuum all the phase velocities and the group velocity are the same. In a dispersive medium, however, they are different because the refractive index is a function of wavelength, which means that the different wavelengths travel at different speeds. Wang and colleagues report evidence for a negative group velocity of -310c, where c (=300 million metres per second) is the speed of light in vacuum."
Tuesday, August 23, 2005
Firing ESB with Synapse
Apache Launches Open Source Software-Integration Project " Apache Synapse would provide many of the capabilities of an enterprise service bus. An ESB, which is available from many vendors, provides secure interoperability between business systems via extensible markup language, web services interfaces and standardized rules-based routing."
The bulk of the project is from Infravio's X-Broker product.
The bulk of the project is from Infravio's X-Broker product.
Monday, August 22, 2005
Things I've Read
* Evangelical Scientists Refute Gravity With New 'Intelligent Falling' Theory related to Boing Boing's $250,000 Intelligent Design challenge (UPDATED: $1 million) and Flying Spaghetti Monster.
* JUnit-Addons and GSBase so at least only two software projects have to reinvent the wheel.
* Non-blocking Data Sharing in Multiprocessor Real-Time Systems "In this paper, we present an efficient non-blocking solution to the general readers/writers inter-task communication problem our solution allows any arbitrary number of readers and writers to perform their respective operations."
* Why I Prefer SOA to REST "The problem is that although it is easy to model resources as services as shown in the example in many cases it is quite difficult to model a service as a resource. For example, a service that validates a credit card number can be modeled as a validateCreditCardNumber(string cardNumber) service. On the other hand it is unintuitive how one would model the service as a resource. For this reason I prefer to think about distributed applications in terms of services as opposed to resources." Now people can just argue about what is intuitive (credit card gateways on the Web existed before SOAP - must have been REST I guess).
* I Need a New Language: Rel? "I want a language for table programming. I think you can write programs in this language that do everything we expect of an application programming language -- building GUIs, reacting to mouse events, listening to sockets -- everything. Don't model your domain as objects. Model it as relations...I can imagine programs that have a relvar (Date's term for a relational variable: essentially a table or a view) for MouseState."
* Package Scoping And Unit Testing "Package scoping particularly shines during unit testing. Some programmers argue that you should only test through the public API. Don't be silly. Limiting your tests to the public API contradicts the spirit of unit testing and subjects you to unnecessary dependency pain. I prefer to isolate and limit the amount of code I test at one time, and test as close to the code as possible." Somewhat related, JSR 277 - Java Module System
* Web as Platform Mash-Ups "There have been a lot of excellent posts and articles this week about APIs, the Web as Platform, web sites as software companies, and so forth..."
* Ruby, Python, "Power" "There are different opinions on the relative power of Ruby and Python. I'm not much more authoritative than other resources (though I'm not less authoritative either; most comparisons between the two languages are flawed). Ultimately I don't believe there are many (any?) places where one language is more "powerful" than the other (and not just in the "they are both Turing complete" sense)"
* JUnit-Addons and GSBase so at least only two software projects have to reinvent the wheel.
* Non-blocking Data Sharing in Multiprocessor Real-Time Systems "In this paper, we present an efficient non-blocking solution to the general readers/writers inter-task communication problem our solution allows any arbitrary number of readers and writers to perform their respective operations."
* Why I Prefer SOA to REST "The problem is that although it is easy to model resources as services as shown in the example in many cases it is quite difficult to model a service as a resource. For example, a service that validates a credit card number can be modeled as a validateCreditCardNumber(string cardNumber) service. On the other hand it is unintuitive how one would model the service as a resource. For this reason I prefer to think about distributed applications in terms of services as opposed to resources." Now people can just argue about what is intuitive (credit card gateways on the Web existed before SOAP - must have been REST I guess).
* I Need a New Language: Rel? "I want a language for table programming. I think you can write programs in this language that do everything we expect of an application programming language -- building GUIs, reacting to mouse events, listening to sockets -- everything. Don't model your domain as objects. Model it as relations...I can imagine programs that have a relvar (Date's term for a relational variable: essentially a table or a view) for MouseState."
* Package Scoping And Unit Testing "Package scoping particularly shines during unit testing. Some programmers argue that you should only test through the public API. Don't be silly. Limiting your tests to the public API contradicts the spirit of unit testing and subjects you to unnecessary dependency pain. I prefer to isolate and limit the amount of code I test at one time, and test as close to the code as possible." Somewhat related, JSR 277 - Java Module System
* Web as Platform Mash-Ups "There have been a lot of excellent posts and articles this week about APIs, the Web as Platform, web sites as software companies, and so forth..."
* Ruby, Python, "Power" "There are different opinions on the relative power of Ruby and Python. I'm not much more authoritative than other resources (though I'm not less authoritative either; most comparisons between the two languages are flawed). Ultimately I don't believe there are many (any?) places where one language is more "powerful" than the other (and not just in the "they are both Turing complete" sense)"
Nullifying C#?
Nullable Types "Support for nullability across all types, including value types, is essential when interacting with databases, yet general purpose programming languages have historically provided little or no support in this area."
Nulls in a database indicate unknown or some other semantic for the value being missing. Missing Information Withot Nulls: "But NULL isn’t a value of type VARCHAR(20), nor of type DECIMAL(6,0)." The decomposition of the tables shown in this looks very much like RDF (page 15 even talks about triple operations).
An example of the syntax:
Nulls not missing anymore "The outcome is that the Nullable type is now a new basic runtime intrinsic. It is still declared as a generic value-type, yet the runtime treats it special. One of the foremost changes is that boxing now honors the null state. A Nullabe int now boxes to become not a boxed Nullable int but a boxed int (or a null reference as the null state may indicate.) Likewise, it is now possible to unbox any kind of boxed value-type into its Nullable type equivalent."
Related: Avoiding Null in C#, Stopping nulls in Java, Null Object Pattern, Null Considered Harmful, Nulls and the Relational Model, Clean design solution for NULLs with Java primitives, Date and Pascal on RDF and Relational Algebra.
While you can argue whether C# should support NULLs in databases, it does seem like a fairly pragmatic decision based on the current state of the industry.
Nulls in a database indicate unknown or some other semantic for the value being missing. Missing Information Withot Nulls: "But NULL isn’t a value of type VARCHAR(20), nor of type DECIMAL(6,0)." The decomposition of the tables shown in this looks very much like RDF (page 15 even talks about triple operations).
An example of the syntax:
int? nFirst = null;
int Second = 2;
nFirst = null; // Valid
Second = nFirst; // Exception, Second is nonnullable.
Nulls not missing anymore "The outcome is that the Nullable type is now a new basic runtime intrinsic. It is still declared as a generic value-type, yet the runtime treats it special. One of the foremost changes is that boxing now honors the null state. A Nullabe int now boxes to become not a boxed Nullable int but a boxed int (or a null reference as the null state may indicate.) Likewise, it is now possible to unbox any kind of boxed value-type into its Nullable type equivalent."
Related: Avoiding Null in C#, Stopping nulls in Java, Null Object Pattern, Null Considered Harmful, Nulls and the Relational Model, Clean design solution for NULLs with Java primitives, Date and Pascal on RDF and Relational Algebra.
While you can argue whether C# should support NULLs in databases, it does seem like a fairly pragmatic decision based on the current state of the industry.
Friday, August 19, 2005
Quick Links
* Songs of the Extremos from the people that brought you The Case Against XP. Maybe in reply someone like Kent Beck will use some Bob Dylan (I wonder who will introduce them). For an interesting view see Bob Dylan: A genius? and The meanings behind the words of the Beatles.
* In trying to find the source of the term "gold plated" in terms of software, I was pointed to "NASA’s Success Checklist": "Implement only what is required. Developers, managers, and customers often think of small, easy changes that seem to make the software better. These changes often have much more far-reaching impacts than anticipated by the specific developer who will implement the change. Do not let additional complexity creep into the project through gold-plating."
* What if VisiCalc had been patented?. Many interesting points, such as: "On the other hand, innovation in VisiCalc-like spreadsheets continued, with Lotus doing things we wouldn't, and then Microsoft moving things further ahead with Excel going in areas Lotus neglected." Also links to, "Patenting VisiCalc".
* Not 2.0? and Web 2.0 or Not?. Such quotes as: "...the key to success in this next stage of the web's evolution is leveraging collective intelligence. And yes, Google's introduction of page rank was absolutely a milestone in this evolution of the web..."
* In trying to find the source of the term "gold plated" in terms of software, I was pointed to "NASA’s Success Checklist": "Implement only what is required. Developers, managers, and customers often think of small, easy changes that seem to make the software better. These changes often have much more far-reaching impacts than anticipated by the specific developer who will implement the change. Do not let additional complexity creep into the project through gold-plating."
* What if VisiCalc had been patented?. Many interesting points, such as: "On the other hand, innovation in VisiCalc-like spreadsheets continued, with Lotus doing things we wouldn't, and then Microsoft moving things further ahead with Excel going in areas Lotus neglected." Also links to, "Patenting VisiCalc".
* Not 2.0? and Web 2.0 or Not?. Such quotes as: "...the key to success in this next stage of the web's evolution is leveraging collective intelligence. And yes, Google's introduction of page rank was absolutely a milestone in this evolution of the web..."
Tuesday, August 16, 2005
Sharper Axe
Sometimes you find some tools that make certain things as easy on Windows as they are on Unix.
StackTrace "Thread dump for Java processes running as a Windows service (like Tomcat, for example), started with javaw.exe or embedded inside another process (Windows and Mac OS X only.)" Also DiffAnywhere.
StackTrace "Thread dump for Java processes running as a Windows service (like Tomcat, for example), started with javaw.exe or embedded inside another process (Windows and Mac OS X only.)" Also DiffAnywhere.
Two Languages Enter, One Language Leaves
Beyond Java "In Beyond Java, Bruce chronicles the rise of the most successful language of all time, and then lays out, in painstaking detail, the compromises the founders had to make to establish success. Then, he describes the characteristics of likely successors to Java. He builds to a rapid and heady climax, presenting alternative languages and frameworks with productivity and innovation unmatched in Java. He closes with an evaluation of the most popular and important programming languages, and their future role in a world beyond Java."
This is a fairly uninformative description.
In a recent interview he said: "I make the point that conditions are ripe for an alternative to emerge. I don't pick what the alternative will be. I just show some of the productivity problems with Java and I show the types of projects in other languages that are interesting. Frameworks like Ruby-on-Rails, Seaside and continuations servers in Lisp and SmallTalk could well become catalysts for another language."
Some more detail is here. It's probable that it's going to mention Ruby. He does mention that: "When you’re mapping a Java class to a schema, you must often type the name of a property five times...Three in the bean: the getter, the setter, the instance variable. One in the schema: the field. Two in the mapping: the property, and the column. In Ruby, you type it once."
But Ruby on Rails is eight times slower than Java. It'll never take off because it's too slow. Where have you heard that before?
This is a fairly uninformative description.
In a recent interview he said: "I make the point that conditions are ripe for an alternative to emerge. I don't pick what the alternative will be. I just show some of the productivity problems with Java and I show the types of projects in other languages that are interesting. Frameworks like Ruby-on-Rails, Seaside and continuations servers in Lisp and SmallTalk could well become catalysts for another language."
Some more detail is here. It's probable that it's going to mention Ruby. He does mention that: "When you’re mapping a Java class to a schema, you must often type the name of a property five times...Three in the bean: the getter, the setter, the instance variable. One in the schema: the field. Two in the mapping: the property, and the column. In Ruby, you type it once."
But Ruby on Rails is eight times slower than Java. It'll never take off because it's too slow. Where have you heard that before?
Monday, August 15, 2005
A World Without Locks
Wikipedia defines lock-free and wait-free algorithms as allowing "...multiple threads to read and write shared data concurrently without corrupting it. "Lock-free" refers to the fact that a thread cannot lock up: every step it takes brings progress to the system."
LOCK-FREE LINKED LISTS AND SKIP LISTS "Developing a correct and efficient memory management scheme is important to make a data structure practical. Developing such a scheme for a lock-free data structure is often quite a challenging task. The difficulty lies in determining how and when memory that was once occupied by parts of the data structure (e.g. nodes of a linked list), can be freed and reused, so that the processes that might still be accessing those parts are able to complete their operations correctly...We presented new algorithms implementing a lock-free linked list and a lock-free skip list. We proved their correctness and lock-freedom."
Lock-Free Reference Counting The goal of this work, therefore, is to allow programmers to exploit the advantages of GC in designing their lock-free data structure implementations, while avoiding its drawbacks. To this end, we provide a methodology that allows programmers to first solve the easier problem of designing a GC-dependent implementation, and to then apply our methodology in order to achieve a GC-independent one.
An older article: Lock-free Parallel Garbage Collection by Mark&Sweep.
Related to Lock Free Programming.
LOCK-FREE LINKED LISTS AND SKIP LISTS "Developing a correct and efficient memory management scheme is important to make a data structure practical. Developing such a scheme for a lock-free data structure is often quite a challenging task. The difficulty lies in determining how and when memory that was once occupied by parts of the data structure (e.g. nodes of a linked list), can be freed and reused, so that the processes that might still be accessing those parts are able to complete their operations correctly...We presented new algorithms implementing a lock-free linked list and a lock-free skip list. We proved their correctness and lock-freedom."
Lock-Free Reference Counting The goal of this work, therefore, is to allow programmers to exploit the advantages of GC in designing their lock-free data structure implementations, while avoiding its drawbacks. To this end, we provide a methodology that allows programmers to first solve the easier problem of designing a GC-dependent implementation, and to then apply our methodology in order to achieve a GC-independent one.
An older article: Lock-free Parallel Garbage Collection by Mark&Sweep.
Related to Lock Free Programming.
Fun in the Sun
* XmlBeansSerializer and Axis 1.3. This looks promising (again). The second in the thread mentions that the XFire project now has server side support. A new startup is now providing support for Axis too. New FLA (four letter acroynm) SASH (Struts, Axis, Spring and Hibernate).
* Introducing AXIOM: The Axis Object Model Introducing AXIOM: The Axis Object Model "AXIOM uses a "builder" that will build the XML object model in memory, according to the events pulled from the underlying StAX parser, but will not create the entire object model at once. Instead, it only builds when the relevant information is absolutely required." Also, OM Tutorial.
* IBM Integrated Ontology Development Toolkit "EODM is the run-time library that allows the application to put in and put out an RDFS/OWL ontology in RDF/XML format; manipulate an ontology using Java objects; call an inference engine and access inference results; and transform between ontology and other models." Also mentions support for SPARQL and OWL DL.
* Encapsulation vs. Inheritance "Inheritance indicates strong encapsulation with other classes, but weak encapsulation between a superclass and its subclasses."
* Introducing AXIOM: The Axis Object Model Introducing AXIOM: The Axis Object Model "AXIOM uses a "builder" that will build the XML object model in memory, according to the events pulled from the underlying StAX parser, but will not create the entire object model at once. Instead, it only builds when the relevant information is absolutely required." Also, OM Tutorial.
* IBM Integrated Ontology Development Toolkit "EODM is the run-time library that allows the application to put in and put out an RDFS/OWL ontology in RDF/XML format; manipulate an ontology using Java objects; call an inference engine and access inference results; and transform between ontology and other models." Also mentions support for SPARQL and OWL DL.
* Encapsulation vs. Inheritance "Inheritance indicates strong encapsulation with other classes, but weak encapsulation between a superclass and its subclasses."
Monday, August 08, 2005
RDF Algebra and Aggregates
* RAL: an Algebra for Querying RDF "RAL is an algebra for RDF defined from a database perspective, some of its operators being inspired by their relational algebra counterparts...Based on the similarities between monads and RAL collections, one can reuse the three monad laws (left unit law, right unit law, and associativity law) as equivalence rules in RAL...RAL operators come in three flavors: extraction operators retrieve the needed resources from the input RDF model, loop operators support repetition, and construction operators build the resulting RDF model." Demonstrates: projection, selection, cartesian product, join, union, difference, intersection, map, Kleene star, create node, create edge, delete node, delete edge, variables and sorting. Part of CognitiveWeb, also has a SPARQL project.
* RDF Aggregate Queries and Views "In this paper, we propose the CAA (Compute Aggregates Algorithm) algorithm to efficiently compute aggregate operations such as COUNT,SUM,AVG,MIN,MAX and so on. CAA can also handle GROUPBY queries. We subsequently define algorithms to maintain aggregate views. These are views involving aggregate queries."
* RDF Aggregate Queries and Views "In this paper, we propose the CAA (Compute Aggregates Algorithm) algorithm to efficiently compute aggregate operations such as COUNT,SUM,AVG,MIN,MAX and so on. CAA can also handle GROUPBY queries. We subsequently define algorithms to maintain aggregate views. These are views involving aggregate queries."
Saturday, August 06, 2005
Speed Racing
* CheckRDFSyntax and Schemarama Revisited "...thinking about our expectations of RDF “validation” can teach us a lot about RDF’s value, about it’s relationship to XML, and about the things we should focus on building next."
* sparql fast as hell "But the real astonishing thing is that SPARQL of this kind is also fast as hell (10-500ms)...In simple words: this is a fulltext scan over all properties of all statements...triplecount: 371994"
* Data First vs. Structure First "Next time you spend energy writing the ontology, or the database schema, or the XML schema, or the software architecture, or the protocol, that 'foresees' problems that you don't have right now think aobut "you ain't gonna need it", "do the simplest thing that can possibly work", "keep it simple stupid", "release early and often", "if ain't broken don't fix it"and all the various other suggestions that tell you not to trust design as the way to solve your problems. But don't forget to think about ways to make further structure emerge from the data, or you'll be lost with a simple system that will fail to grow in complexity without deteriorating." Rifting on a familar theme.
* PlayStation 3 processor could support Mac OS X Tiger ""The operating system has also yet to be clarified. The integrated Cell processor will be able to support a variety of operating systems (such as Linux or Apple's Tiger)."
* CollectionClosureMethod "The each method takes a one argument block (Ruby and Smalltalk both refer to closures as blocks). It then executes the block on each element in the collection. It essentially is the same as the foreach statement you find in many modern languages (and recently arrived in Java with 1.5). With these languages the foreach method is all you get, but with collections and closures the each method is just the start."
* sparql fast as hell "But the real astonishing thing is that SPARQL of this kind is also fast as hell (10-500ms)...In simple words: this is a fulltext scan over all properties of all statements...triplecount: 371994"
* Data First vs. Structure First "Next time you spend energy writing the ontology, or the database schema, or the XML schema, or the software architecture, or the protocol, that 'foresees' problems that you don't have right now think aobut "you ain't gonna need it", "do the simplest thing that can possibly work", "keep it simple stupid", "release early and often", "if ain't broken don't fix it"and all the various other suggestions that tell you not to trust design as the way to solve your problems. But don't forget to think about ways to make further structure emerge from the data, or you'll be lost with a simple system that will fail to grow in complexity without deteriorating." Rifting on a familar theme.
* PlayStation 3 processor could support Mac OS X Tiger ""The operating system has also yet to be clarified. The integrated Cell processor will be able to support a variety of operating systems (such as Linux or Apple's Tiger)."
* CollectionClosureMethod "The each method takes a one argument block (Ruby and Smalltalk both refer to closures as blocks). It then executes the block on each element in the collection. It essentially is the same as the foreach statement you find in many modern languages (and recently arrived in Java with 1.5). With these languages the foreach method is all you get, but with collections and closures the each method is just the start."
Wednesday, August 03, 2005
Duplication is a Mistake
Again, I'm looking at DISTINCT in SPARQL.
In the relational world Date talks about how users don't care about duplicates and it makes optimization difficult and invalidates operations (like JOIN). Preventing duplicates means that optimizers can make logically equivalent transformations. It seems quite valid to made distinct results the only option.
An example he gives is a query to get supplier numbers for suppliers who supply at least one part , "DOUBLE TROUBLE, DOUBLE TROUBLE PART 1": "The obvious first point to make is that the twelve different formulations produce nine different results! -- different, that is, with respect to their degree of duplication...Thus, if the user really cares about duplicates, then he or she needs to be extremely careful in formulating the query appropriately in order to obtain exactly the desired result."
"Here are some implications of this point:
* First, the optimizer code itself is harder to write, harder to maintain, and probably more buggy--all of which combines to make the product simultaneously more expensive and less reliable, as well as late in delivery in the marketplace.
* Second, system performance is likely to be worse than it might otherwise be.
* Third, the user is going to have to get involved in performance issues; for instance, the user might have to spend time and effort on figuring out the best way to express a given query (a state of affairs, incidentally, that the relational model was explicitly designed to avoid)."
"...if I say "the sun is shining here today" and "the sun is shining here today," I'm simply telling you the sun is shining here today! And from this perspective, the notion of duplicate rows--as that notion is usually understood--obviously makes no sense at all."
There's also a part two.
The same point is made here: "I think it would be a mistake for the query language to take a position on whether or not query result sets could contain duplicate rows (or if it did take a position, I'd want it to be that they couldn't!) From a selfish perspective, I worry that we'll have to de-tune RDF Gateway's query evaluation in order to allow duplicate rows to exist in a resultset (after all if a user wants duplicate rows, they can merely select out the variable(s) that make those rows distinguishable). Perhaps the issue of duplicate rows could be implementation specific?"
It seems that Danny is reading the same thing I am.
In the relational world Date talks about how users don't care about duplicates and it makes optimization difficult and invalidates operations (like JOIN). Preventing duplicates means that optimizers can make logically equivalent transformations. It seems quite valid to made distinct results the only option.
An example he gives is a query to get supplier numbers for suppliers who supply at least one part , "DOUBLE TROUBLE, DOUBLE TROUBLE PART 1": "The obvious first point to make is that the twelve different formulations produce nine different results! -- different, that is, with respect to their degree of duplication...Thus, if the user really cares about duplicates, then he or she needs to be extremely careful in formulating the query appropriately in order to obtain exactly the desired result."
"Here are some implications of this point:
* First, the optimizer code itself is harder to write, harder to maintain, and probably more buggy--all of which combines to make the product simultaneously more expensive and less reliable, as well as late in delivery in the marketplace.
* Second, system performance is likely to be worse than it might otherwise be.
* Third, the user is going to have to get involved in performance issues; for instance, the user might have to spend time and effort on figuring out the best way to express a given query (a state of affairs, incidentally, that the relational model was explicitly designed to avoid)."
"...if I say "the sun is shining here today" and "the sun is shining here today," I'm simply telling you the sun is shining here today! And from this perspective, the notion of duplicate rows--as that notion is usually understood--obviously makes no sense at all."
There's also a part two.
The same point is made here: "I think it would be a mistake for the query language to take a position on whether or not query result sets could contain duplicate rows (or if it did take a position, I'd want it to be that they couldn't!) From a selfish perspective, I worry that we'll have to de-tune RDF Gateway's query evaluation in order to allow duplicate rows to exist in a resultset (after all if a user wants duplicate rows, they can merely select out the variable(s) that make those rows distinguishable). Perhaps the issue of duplicate rows could be implementation specific?"
It seems that Danny is reading the same thing I am.
Monday, August 01, 2005
VFS
VFS " Commons VFS provides a single API for accessing various different file systems. It presents a uniform view of the files from various different sources, such as the files on local disk, on an HTTP server, or inside a Zip archive."
Also supports WebDAV and CIFS (Samba). Comes with a file system abstraction that includes junctions (links) and listeners.
Also supports WebDAV and CIFS (Samba). Comes with a file system abstraction that includes junctions (links) and listeners.
Thursday, July 28, 2005
Enough of Enums
Another one of these by type rather than instanceof or if-else. Joshua Bloch seems to have the answer: "If a typesafe enum class has methods whose behavior varies significantly from one class constant to another, you should use a separate private nested class or anonymous inner class for each constant. This allows each constant to have its own implementation of each such method, and automatically invokes the correct implementation. The alternative is to structure each such method as a multi-way branch that behaves differently depending on the constant on which it's invoked. This alternative is ugly, error prone, and likely to provide performance that is inferior to that of the virtual machine's automatic method dispatching."
Object Input Classes "Note – The readResolve method is not invoked on the object until the object is fully constructed, so any references to this object in its object graph will not be updated to the new object nominated by readResolve. However, during the serialization of an object with the writeReplace method, all references to the original object in the replacement object’s object graph are replaced with references to the replacement object. Therefore in cases where an object being serialized nominates a replacement object whose object graph has a reference to the original object, deserialization will result in an incorrect graph of objects. Furthermore, if the reference types of the object being read (nominated by writeReplace) and the original object are not compatible, the construction of the object graph will raise a ClassCastException."
Exploring Enums: The Wait Is Finally Over "Spare the serialization: Enum serialization isn't like the normal one you have seen. The process by which enum constants are serialized cannot be customized. Any class-specific writeObject and writeReplace methods defined by enum types are ignored during serialization. Similarly, any serialPersistentFields or serialVersionUID field declarations are also ignored - all enum types have a fixed serialVersionUID of 0L. Again, this shouldn't concern you too much. Let the language take care of the specifics."
Object Input Classes "Note – The readResolve method is not invoked on the object until the object is fully constructed, so any references to this object in its object graph will not be updated to the new object nominated by readResolve. However, during the serialization of an object with the writeReplace method, all references to the original object in the replacement object’s object graph are replaced with references to the replacement object. Therefore in cases where an object being serialized nominates a replacement object whose object graph has a reference to the original object, deserialization will result in an incorrect graph of objects. Furthermore, if the reference types of the object being read (nominated by writeReplace) and the original object are not compatible, the construction of the object graph will raise a ClassCastException."
Exploring Enums: The Wait Is Finally Over "Spare the serialization: Enum serialization isn't like the normal one you have seen. The process by which enum constants are serialized cannot be customized. Any class-specific writeObject and writeReplace methods defined by enum types are ignored during serialization. Similarly, any serialPersistentFields or serialVersionUID field declarations are also ignored - all enum types have a fixed serialVersionUID of 0L. Again, this shouldn't concern you too much. Let the language take care of the specifics."
Wednesday, July 27, 2005
Pre-Conditions and Post-Conditions
Rules based routing "The basic idea is you expose a DroolsComponent at some service/interface/operation endpoint in ServiceMix then let it perform rules based routing, or other actions as required.
You can deploy a DroolsComponent with a rule base which will be fired when it is invoked. The rule base is then in complete control over messge dispatching."
JBI Routing "...a component can give the container some hints by specifying the service, interface and/or operation to invoke and then let the container choose which physical service endpoint to invoke using some choosing algorithm (maybe using rules or policy driven metadata). Remember there may be many services available for a specific service, interface and operation names."
This is almost the exact same idea that I had on a previous project. The routing of a document was done based on it's metadata as well as system being configured for specific outcomes. For example, you could configure it such that all documents should be turned into HTML for example.
Via Rules based routing in JBI using Drools in ServiceMix.
You can deploy a DroolsComponent with a rule base which will be fired when it is invoked. The rule base is then in complete control over messge dispatching."
JBI Routing "...a component can give the container some hints by specifying the service, interface and/or operation to invoke and then let the container choose which physical service endpoint to invoke using some choosing algorithm (maybe using rules or policy driven metadata). Remember there may be many services available for a specific service, interface and operation names."
This is almost the exact same idea that I had on a previous project. The routing of a document was done based on it's metadata as well as system being configured for specific outcomes. For example, you could configure it such that all documents should be turned into HTML for example.
Via Rules based routing in JBI using Drools in ServiceMix.
Tuesday, July 26, 2005
Whuffie and Defeasible Inference
Whuffie "Many community-oriented websites are experimenting with Whuffie-like concepts of reputation management (Slashdot's karma system, for example, or eBay's feedback ratings). Others look further ahead, at the "next generation" of the web - known as the Semantic Web.
One of the key challenges in developing the Semantic Web is in fact Whuffie, although you won't hear it called by that name.
Built of assertations about facts, the Semantic Web is basically distributed metadata. If one party says "Water is Wet" while another claims "Water isn't Wet", problems are encountered. At this point, Whuffie plays a role: which party is trusted more by others whom I already trust?
One of the key researchers in this area is Jennifer Golbeck, who is performing research on Trust in the Semantic Web"
This seems to imply that the Semantic Web offers non-monotonic, defeasible inference. Maybe it's just not a very clear example. A good run down is Non-Monotonic Logic (Nixon and Tweety are given examples, but not Clyde).
One of the key challenges in developing the Semantic Web is in fact Whuffie, although you won't hear it called by that name.
Built of assertations about facts, the Semantic Web is basically distributed metadata. If one party says "Water is Wet" while another claims "Water isn't Wet", problems are encountered. At this point, Whuffie plays a role: which party is trusted more by others whom I already trust?
One of the key researchers in this area is Jennifer Golbeck, who is performing research on Trust in the Semantic Web"
This seems to imply that the Semantic Web offers non-monotonic, defeasible inference. Maybe it's just not a very clear example. A good run down is Non-Monotonic Logic (Nixon and Tweety are given examples, but not Clyde).
Esquilax
Dynamic Typing in C# "The effect is pseudo-dynamic typing—dynamic style programming on top of a statically typed language. The introduction of pseudo-dynamic typing in C# raises the possibility that it may become pervasive as a scripting language, stealing much thunder away from Python, Perl and Ruby."
Links to: Static Typing Where Possible, Dynamic Typing When Needed: The End of the Cold War Between Programming Languages.
Links to: Static Typing Where Possible, Dynamic Typing When Needed: The End of the Cold War Between Programming Languages.
APIs on REST
* XINS is a technology used to define, create and invoke remote APIs. XINS is specification-oriented. When API specifications are written (in XML), XINS will transform them to HTML-based documentation and Java code for both the client- and the server-side. The communication is based on HTTP.
* Axis 2.0 Support for REST "Axis2 can be configured as REST Cantainer and can be used to send and receive restful web services requests and responses. The REST Web Services can be access in two ways, using HTTP GET and POST."
* REST vs API "The starting point is usually that somebody has an API that is intended to shield the developer from the inner workings of SOAP and perhaps another protocol or three. The person is thinking about adding REST support (generally in the form of removing the requirement for a SOAP envelope and adding support for additional HTTP methods). What can go wrong?"
* Axis 2.0 Support for REST "Axis2 can be configured as REST Cantainer and can be used to send and receive restful web services requests and responses. The REST Web Services can be access in two ways, using HTTP GET and POST."
* REST vs API "The starting point is usually that somebody has an API that is intended to shield the developer from the inner workings of SOAP and perhaps another protocol or three. The person is thinking about adding REST support (generally in the form of removing the requirement for a SOAP envelope and adding support for additional HTTP methods). What can go wrong?"
Monday, July 18, 2005
Pasta Preference
Ravioli code "The opposite of spaghetti code, where too many small objects that rely on many other small objects are created."
See also spaghetti code.
See also spaghetti code.
Thursday, July 14, 2005
Intel Macs
Speed of Apple Intel dev systems impress developers "Developers are renting the $999 hardware from Apple for a period of 18 months in order to get a head start in porting their applications to run on the Intel version of Mac OS X.
"It's fast," said one developer source of Mac OS X running on Intel's Pentium processors. "Faster than [Mac OS X] on my Dual 2GHz Power Mac G5." In addition to booting Windows XP at blazing speeds, the included version of Mac OS X for Intel takes "as little as 10 seconds" to boot to the Desktop from when the Apple logo first displays on screen." Via Slashdot. Transition kit is available here.
WebObjects 5.3 Which is part of XCode, which is part of the OS.
"It's fast," said one developer source of Mac OS X running on Intel's Pentium processors. "Faster than [Mac OS X] on my Dual 2GHz Power Mac G5." In addition to booting Windows XP at blazing speeds, the included version of Mac OS X for Intel takes "as little as 10 seconds" to boot to the Desktop from when the Apple logo first displays on screen." Via Slashdot. Transition kit is available here.
WebObjects 5.3 Which is part of XCode, which is part of the OS.
Wednesday, July 13, 2005
Less Code
Started the first go of using generics in JRDF. It uses Java's collections pretty heavily so there's quite a lot of work to be done. Many good things happened, mostly it's less code - casting went away (of course) and things like add became type safe. It makes a lot more sense to add/remove an org.jrdf.graph.ObjectNode to a Bag rather than an java.lang.Object.
One thing I have yet to do (and see if it's even possible) is to see whether things like addAll, containsAll, etc can be made to use more specific variants (Alternative, Bag, etc) rather than Collection. From what I've gathered so far it may not be possible for backwards compatibility reasons. It seems like what I want (and what Java generics needs apparently) is variant types.
Tom has some comments on 0.3.4 too.
Update: In some instances (pun intended) there's more code: How do I generically create objects and arrays? and How do I perform a runtime type check whose target type is a type parameter? "For a type check at runtime we must explicitly provide runtime type information so that we can perform the type check and cast by means of reflection. The type information is best supplied by means of a Class object."
One thing I have yet to do (and see if it's even possible) is to see whether things like addAll, containsAll, etc can be made to use more specific variants (Alternative, Bag, etc) rather than Collection. From what I've gathered so far it may not be possible for backwards compatibility reasons. It seems like what I want (and what Java generics needs apparently) is variant types.
Tom has some comments on 0.3.4 too.
Update: In some instances (pun intended) there's more code: How do I generically create objects and arrays? and How do I perform a runtime type check whose target type is a type parameter? "For a type check at runtime we must explicitly provide runtime type information so that we can perform the type check and cast by means of reflection. The type information is best supplied by means of a Class object."
JRDF 0.3.4
JRDF Comes with the port of Sesame's RIO parser to JRDF, the start of a SPARQL parser and remote querying, bug fixes and other small changes. This is the last version for Java 1.4. The next version will use some of the new 5.0 features like Generics for RDF collections (seems to make sense) and the much touted concurrency libraries.
Tuesday, July 12, 2005
Oracle Releases Database with RDF Support
Oracle supports RDF Danny points to the fact that Oracle 10g now supports.
First announced about a year ago.
First announced about a year ago.
Monday, July 11, 2005
Objectifying the Web
REST versus Object-Orientation (and a little python) "From my perspective, the concepts of Object Orientation (OO) and REST are comparable. They both seek to identify "things" that correspond to something you can about and interact with as a unit. In OO we call it an object. In REST we call it a resource. Both OO and REST allow some form of abstraction. You can replace one object with another of the same class or of the same type without changing how client code interacts with the object. The type is a common abstraction that represents any suitable object equally well."
REST and Object Orientation via OWL "So the most important aspect of RESTful programming as Benjamin points out is that now all objects are universally nameable and universally accessible. A few simple verbs allow us to create, delete and change the state of these resources. Access Control Lists allow fine tuning of responsibilities for a resource. And OWL gives us the Object Oriented conceptual structure to predict and understand the content of these resources and how they relate to others."
People seem to forget that method calls in Java or whatever can be considered message passing. What could be more OO than REST? GET isn't analogous to "getFoo" or "getBar" on a JavaBean object. The semantics behind GET/POST/PUT/DELETE is better thought of as how you are changing the object (related to Create Retrieve Update Delete (CRUD)). What could be more RESTful than OO?
The Dining Philosophers in REST "...it's trivial to encapsulate transactions into a single atomic exchange, and exposing those encapsulations as first class entities is generally a pretty good idea. And it certainly fits with the REST model of the world. MUCH more so, in fact, than RPC, as it forces one to aggregate those operations in the server design rather than leaving it open to the client to abuse. In fact, I'd even say that doing so is critical to the ability to build workable distributed systems."
Also related "Using RDF to improve Object-First Development".
REST and Object Orientation via OWL "So the most important aspect of RESTful programming as Benjamin points out is that now all objects are universally nameable and universally accessible. A few simple verbs allow us to create, delete and change the state of these resources. Access Control Lists allow fine tuning of responsibilities for a resource. And OWL gives us the Object Oriented conceptual structure to predict and understand the content of these resources and how they relate to others."
People seem to forget that method calls in Java or whatever can be considered message passing. What could be more OO than REST? GET isn't analogous to "getFoo" or "getBar" on a JavaBean object. The semantics behind GET/POST/PUT/DELETE is better thought of as how you are changing the object (related to Create Retrieve Update Delete (CRUD)). What could be more RESTful than OO?
The Dining Philosophers in REST "...it's trivial to encapsulate transactions into a single atomic exchange, and exposing those encapsulations as first class entities is generally a pretty good idea. And it certainly fits with the REST model of the world. MUCH more so, in fact, than RPC, as it forces one to aggregate those operations in the server design rather than leaving it open to the client to abuse. In fact, I'd even say that doing so is critical to the ability to build workable distributed systems."
Also related "Using RDF to improve Object-First Development".
Free Expression of Ideas
* Even if you think you can beat Waldo, you can't beat Einstein "For example to ping a peer on my local wireless network the best latency (as measured by running batch of pings) is around 2ms but my average latency is 100ms over a negligible distance compared to the speed of light. In the average case on this network the binary protocol is a mere 14% faster than SOAP - certainly not tens of times faster."
* Speed of Light Fallacy "If you consider these numbers, you will see that the speed-of-light argument is a fallacy. Dealing with the speed-of-light limitation is achieved by minimizing the number of messages that are exchanged, and by maximizing the amount of data sent with each message, regardless of the transfer technology."
* Generics Considered Harmful "The complexity of Java has been turbocharged to what seems to me relatively small benefit. I don't see that the value is there to justify the cost. Not that we can change things, but I think we should at least view it as an demonstration proof of the value of an explicit complexity budget against which features must be justified. Without such a budget, it feels like the JSR process ran far ahead, without a step back to ask “Is this feature really necessary”. It seemed to just be understood that it was necessary."
* Generics Considered Harmful? Nope. "For a few library designers, who now want to create generic types (they still have the option of creating old style types), life just got more complicated. However, for all their clients, life got easier. Easier is good."
* Speed of Light Fallacy "If you consider these numbers, you will see that the speed-of-light argument is a fallacy. Dealing with the speed-of-light limitation is achieved by minimizing the number of messages that are exchanged, and by maximizing the amount of data sent with each message, regardless of the transfer technology."
* Generics Considered Harmful "The complexity of Java has been turbocharged to what seems to me relatively small benefit. I don't see that the value is there to justify the cost. Not that we can change things, but I think we should at least view it as an demonstration proof of the value of an explicit complexity budget against which features must be justified. Without such a budget, it feels like the JSR process ran far ahead, without a step back to ask “Is this feature really necessary”. It seemed to just be understood that it was necessary."
* Generics Considered Harmful? Nope. "For a few library designers, who now want to create generic types (they still have the option of creating old style types), life just got more complicated. However, for all their clients, life got easier. Easier is good."
Sunday, July 10, 2005
Visual Representation of Build Status
Automated Continuous Integration and the Ambient Orb™ Forget manually controlled Lava Lamps (as currently used at work) or plain light bulbs (note the use of 3 colours and semantics for various "on" positions).
Via Other UIs and Google searching (because it appeared so obvious).
Programmable via a serial port. More information is available on the developer page.
Via Other UIs and Google searching (because it appeared so obvious).
Programmable via a serial port. More information is available on the developer page.
Querying for Hierachies
Storing Hierarchical Data in a Database "Whether you want to build your own forum, publish the messages from a mailing list on your Website, or write your own cms: there will be a moment that you’ll want to store hierarchical data in a database. And, unless you’re using a XML-like database, tables aren’t hierarchical; they’re just a flat list. You’ll have to find a way to translate the hierarchy in a flat file."
"If you want to display the tree using a table with left and right values, you’ll first have to identify the nodes that you want to retrieve. For example, if you want the ‘Fruit’ subtree, you’ll have to select only the nodes with a left value between 2 and 11. In SQL, that would be:
SELECT * FROM tree WHERE lft BETWEEN 2 AND 11;"
It struct me how complicated either of these solutions are.
In iTQL it would be using walk:
"select $subject
...
where walk($subject <rdfs:subClassOf> <food:fruit>
and $subject <rdfs:subClassOf> $object);
Or in Sesame's SeRQL (concrete example here):
"SELECT DISTINCT _fruit
FROM _fruit serql:directSubClassOf <food:fruit>"
They aren't quite the same query - Sesame is inferring new statements and walk is not. In designing iTQL we were always deciding between magical predicates or functions (like walk) and sometimes slavishly sticking to triple patterns. Combining the walk and trans operations using a predicate, like in SeRQL, seems like a good approach to take.
Also, interesting in a fairly recent Sesame release are the ANY and ALL, EXISTS and MINUS features.
"If you want to display the tree using a table with left and right values, you’ll first have to identify the nodes that you want to retrieve. For example, if you want the ‘Fruit’ subtree, you’ll have to select only the nodes with a left value between 2 and 11. In SQL, that would be:
SELECT * FROM tree WHERE lft BETWEEN 2 AND 11;"
It struct me how complicated either of these solutions are.
In iTQL it would be using walk:
"select $subject
...
where walk($subject <rdfs:subClassOf> <food:fruit>
and $subject <rdfs:subClassOf> $object);
Or in Sesame's SeRQL (concrete example here):
"SELECT DISTINCT _fruit
FROM _fruit serql:directSubClassOf <food:fruit>"
They aren't quite the same query - Sesame is inferring new statements and walk is not. In designing iTQL we were always deciding between magical predicates or functions (like walk) and sometimes slavishly sticking to triple patterns. Combining the walk and trans operations using a predicate, like in SeRQL, seems like a good approach to take.
Also, interesting in a fairly recent Sesame release are the ANY and ALL, EXISTS and MINUS features.
Tuesday, July 05, 2005
The Singleton and Double Checked Locking
The "Double-Checked Locking is Broken" Declaration "If the singleton you are creating is static (i.e., there will only be one Helper created), as opposed to a property of another object (e.g., there will be one Helper for each Foo object, there is a simple and elegant solution.
Just define the singleton as a static field in a separate class. The semantics of Java guarantee that the field will not be initialized until the field is referenced, and that any thread which accesses the field will see all of the writes resulting from initializing that field.
class HelperSingleton {
static Helper singleton = new Helper();
}"
I've had this happen in real life many times, where people expect DCL to work. It doesn't no matter how hard you try (and I've had seen some really smart people try).
There is no way around it except for Java 1.5.
With something like Spring and PicoContainer, though, the Singleton pattern is fairly inherent - a configuration option for an object rather than something you specify in code.
Just define the singleton as a static field in a separate class. The semantics of Java guarantee that the field will not be initialized until the field is referenced, and that any thread which accesses the field will see all of the writes resulting from initializing that field.
class HelperSingleton {
static Helper singleton = new Helper();
}"
I've had this happen in real life many times, where people expect DCL to work. It doesn't no matter how hard you try (and I've had seen some really smart people try).
There is no way around it except for Java 1.5.
With something like Spring and PicoContainer, though, the Singleton pattern is fairly inherent - a configuration option for an object rather than something you specify in code.
Apple eXtending RSS
RSS inventor slams Apple's podcast approach "Extending a standard is not, in and of itself, a bad idea. The "X" in "XML" does continue to stand for "Extensible." Conceivably, Apple could have made an open proposition for an extension to media RSS that supports the iTunes catalog, stated Louis, and RSS developers could have worked with them to polish a standard prior to iTunes 4.9's release. But by redefining the meaning of at least two principal RSS tags, by Winer's count, Apple may be, in Louis' opinion, trying too hard to re-invent the wheel just so the final product can be theirs and theirs alone. The situation is reminiscent, said Louis, of the browser wars of the 1990s "where basically you could develop your own tags that could work on your own browser, but then there was no guarantee that the community would support those.""
Friday, July 01, 2005
Cringley on Web Two Point Oh
Accessories Make the Nerd "...the gray web is filled with data that we can search, perhaps, but can't understand. Imagine using an English-language search engine to search a Persian-language web site. The way out of this, to a new dawn where visible, invisible, and gray data alike are available to us, is through Web 2.0 (sometimes called or confused with the so-called "semantic web"), where we will use metadata (primarily XML) to advertise our needs and disposals to the world.
This is a huge leap for anything as established as the Web, to say that we are going to add an overlay of metadata so that you can not only share with the world a picture of your car, but also set terms under which you'd sell that car."
"There are several problems with this concept, not the least of which is the Tower of Babel effect in which every metadata tagger can use his or her own tagging system, none of which are necessarily readable by the others. Joe sees no problem in that because computers are really good at lifting and carrying, and mapping a thousand tag formats to each other isn't all that much harder than mapping a few. But I see people as being inherently lazy, and Yahoo and Google and all the other big web companies as being inherently greedy and unwilling to give up their businesses so easily."
"Here is what Web 2.0 WILL be, in my view: a new way of structuring Internet businesses around published APIs, Application Programming Interfaces. New companies will spring up that simply glue web-based APIs together...Web 2.0 will be staffed by two different kinds of entrepreneurs -- those who provide staunch web services exposed through APIs (Amazon, eBay, Google, and a bunch more), and those who glue those services together and make some sort of useful abstraction service."
This is a huge leap for anything as established as the Web, to say that we are going to add an overlay of metadata so that you can not only share with the world a picture of your car, but also set terms under which you'd sell that car."
"There are several problems with this concept, not the least of which is the Tower of Babel effect in which every metadata tagger can use his or her own tagging system, none of which are necessarily readable by the others. Joe sees no problem in that because computers are really good at lifting and carrying, and mapping a thousand tag formats to each other isn't all that much harder than mapping a few. But I see people as being inherently lazy, and Yahoo and Google and all the other big web companies as being inherently greedy and unwilling to give up their businesses so easily."
"Here is what Web 2.0 WILL be, in my view: a new way of structuring Internet businesses around published APIs, Application Programming Interfaces. New companies will spring up that simply glue web-based APIs together...Web 2.0 will be staffed by two different kinds of entrepreneurs -- those who provide staunch web services exposed through APIs (Amazon, eBay, Google, and a bunch more), and those who glue those services together and make some sort of useful abstraction service."
Thursday, June 30, 2005
Semantic Web Fast, SOAP not that Slow and other links
* The Semantic Web In One Day "...syntactic aspects of data integration turned out to be tedious. Often, output from tool A can’t be used directly as input for tool B, although both have the same language capabilities. For example, both tools can handle RDF for input and output, but the resulting data is syntactically incompatible to the extent that the tools can’t communicate." Full article here.
* SOAP Performance Considered Really Rather Good points to a number of people studying the speed of SOAP. An intersting paper is "An Evaluation of Contemporary Commercial SOAP Implementations" which says that "SOAP and non-SOAP implementations continued to widen with .NET Remoting offering 280 msgs/sec at peak while most SOAP implementations were only handling from 30 to 60 msgs/sec. Even the leading Product A Document/Literal implementation only gave a maximum throughput of 67 msgs/sec. The two lowest performing RPC/Encoded implementations only handled 15 msgs/sec, the binary/TCP alternative." Not quite the "speed of light is the limiting factor".
* Secrets of the A-list bloggers: Technorati vs. Google "If Google favors indexing more popular sites more often, a clear opprtunity for world-live-web search engines like Technorati would be in the long tail of less-often-indexed sites but Technorati seems to ignore that opportunity and concentrate on the top sites. What that will translate into is a direct reproduction of the power laws when it comes to indexing of blogs."
* A conversation with Jeff Nielsen about agile software development "I was particularly interested to hear about Jeff's use of FIT, Ward Cunningham's Framework for Integrated Test. This technique first appeared on my radar in an outtake from our 2003 story on test-driven development. A more recent development is Fitnesse, a Wiki that supports the use of FIT... pains me to say so but, according to Jeff, XML-oriented tools have so far failed to cut the mustard in this environment." XML is not agile!
* Managing Component Dependencies Using ClassLoaders "Java's class loading mechanism allows for more elegant solutions to this problem. One such solution is for each component's authors to specify the dependencies of their component inside of its JAR manifest."
* SOAP Performance Considered Really Rather Good points to a number of people studying the speed of SOAP. An intersting paper is "An Evaluation of Contemporary Commercial SOAP Implementations" which says that "SOAP and non-SOAP implementations continued to widen with .NET Remoting offering 280 msgs/sec at peak while most SOAP implementations were only handling from 30 to 60 msgs/sec. Even the leading Product A Document/Literal implementation only gave a maximum throughput of 67 msgs/sec. The two lowest performing RPC/Encoded implementations only handled 15 msgs/sec, the binary/TCP alternative." Not quite the "speed of light is the limiting factor".
* Secrets of the A-list bloggers: Technorati vs. Google "If Google favors indexing more popular sites more often, a clear opprtunity for world-live-web search engines like Technorati would be in the long tail of less-often-indexed sites but Technorati seems to ignore that opportunity and concentrate on the top sites. What that will translate into is a direct reproduction of the power laws when it comes to indexing of blogs."
* A conversation with Jeff Nielsen about agile software development "I was particularly interested to hear about Jeff's use of FIT, Ward Cunningham's Framework for Integrated Test. This technique first appeared on my radar in an outtake from our 2003 story on test-driven development. A more recent development is Fitnesse, a Wiki that supports the use of FIT... pains me to say so but, according to Jeff, XML-oriented tools have so far failed to cut the mustard in this environment." XML is not agile!
* Managing Component Dependencies Using ClassLoaders "Java's class loading mechanism allows for more elegant solutions to this problem. One such solution is for each component's authors to specify the dependencies of their component inside of its JAR manifest."
It's somehow fuzzy here
Playing with Google Earth the US area has fast food places and monuments and the rest of the world is all fuzzy and you most just get rivers drawn up. You can tell where the closest KFC is to the Washington monument but I don't even know if there's a Krispy Kreme Doughnuts anywhere in Australia. Is this a realistic view of the rest of the world from an American point of view? I haven't checked but is Iraq highly detailed, Afghanistan all blurry (especially around the borders), Canada is just like the US, etc.
Tuesday, June 28, 2005
Northrop Buys Tucana and Continues Kowari
Northrop Grumman Buys the Tucana Knowledge Server "Stunningly for a company their size, Northrop has not only agreed to support Kowari but rushed to do so. I certainly didn't expect a US federal systems integrator to "get" Open Source Software, but times have clearly changed. Their senior managers have made a legitimate effort to figure out the licensing and how to make it work within their business model. I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
A Grumman employee? Or a Thoughtworker with inside knowledge? Building an Agile Enterprise "I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
Update: Paul's doing the handover
A Grumman employee? Or a Thoughtworker with inside knowledge? Building an Agile Enterprise "I have confidence that we can figure out a way to make it work for both the Kowari community and Northrop Grumman."
Update: Paul's doing the handover
Wednesday, June 22, 2005
Jazzed by Jackrabbit
Catch Jackrabbit and the Java Content Repository API "If the Java Content Repository (JCR) API expert group's vision bears out, in five or ten years' time we will all program to repositories, not databases, according to David Nuescheler, CTO of Day Software [4], and JSR 170 spec lead. Repositories are an outgrowth of many years of data management research, and are best understood as fancy object stores especially suited to today's applications."
"The Jackrabbit code base contains not only the JCR API reference implementation, but also a fully functional repository as well as several contributed libraries for tasks, such as accessing a remote repository via RMI. There is even a JDBC persistence manager to allow plugging in a relational database as a persistent store, and an object-relational mapping tool that allows Hibernate applications to use the repository."
"The default Jackrabbit repository is based on the file system. However, Jackrabbit provides a JDBC persistence manager that relegates data storage to a relational database. As any JCR-compliant repository, Jackrabbit can be accessed through any protocol such as WebDAV or RMI. Examples for different repository access modes are included in the Jackrabbit source distribution."
"Still, for a public blogging "superstore" to have real value to application developers, for instance, some agreement on the node types supporting a blogging data model would be helpful. The repository community has so far avoided the politically sensitive pitfalls of trying to initiate agreement about such information models. The jury is still out whether truly universal data "superstores" can emerge in the absence of such shared data models, or if they will remain a dream befitting Utopia."
"The Jackrabbit code base contains not only the JCR API reference implementation, but also a fully functional repository as well as several contributed libraries for tasks, such as accessing a remote repository via RMI. There is even a JDBC persistence manager to allow plugging in a relational database as a persistent store, and an object-relational mapping tool that allows Hibernate applications to use the repository."
"The default Jackrabbit repository is based on the file system. However, Jackrabbit provides a JDBC persistence manager that relegates data storage to a relational database. As any JCR-compliant repository, Jackrabbit can be accessed through any protocol such as WebDAV or RMI. Examples for different repository access modes are included in the Jackrabbit source distribution."
"Still, for a public blogging "superstore" to have real value to application developers, for instance, some agreement on the node types supporting a blogging data model would be helpful. The repository community has so far avoided the politically sensitive pitfalls of trying to initiate agreement about such information models. The jury is still out whether truly universal data "superstores" can emerge in the absence of such shared data models, or if they will remain a dream befitting Utopia."
Tuesday, June 21, 2005
Scripting the Semantic Web
Semantic Web, meet Ruby on Rails "OWL is capable of defining a rich object model with classes, properties and instances, which led me to consider the possibility of a separating domain model maintenance from the implementation of services that leverage the model. I've seen attempts to realize this vision, but the only ones I can remember involve code generation, particularly from UML models and have not been particularly successful. Ruby, being a dynamic language with rich metaprogramming facilities presents some interesting possibilities for using shared domain models since it allows runtime definition of classes. It should be possible to write a library for Ruby that loads entire object models and instances on the fly, much in the same manner that Ruby on Rails ActiveRecord loads its class definitions for persistent classes directly from database metadata."
Also related, Deep Integration of Scripting Languages and Semantic Web Technologies "Instead of loading an ontology, the notion of importing one leads to a different association for the programmer. The difference may be regarded as pedantic or subtle, but it is an important one: the programmer, instead of regarding the ontology as mere data she has to load and access via an API, the imported ontology behaves like a library, extending her possibilities like only code does."
What's covered in the second paper is the problem of the Open World Assumption, in programming they state that we need a third value "unknown" rather than just true and false.
Somewhat related, Hitting reload is the framework job "What's the point in designing tables for a webapp when an RDF-backed store will manage the data for you and RDF queries will come back as tabular data anyway? There are RDF triple stores that will handle in the order 10^6 statements - Leigh Dodds is doing some research on that, up to 10^8 by the looks of things. If I need queries instead of hacking out iterators+fiters I'll use versa/itql/rdql. Now, saying I never want to design another relational schema again is not to say I don't want to use a database. Most of these RDF triple stores are in fact using an RDBMS in the background, as the filesystem and indexer, it's just that the relational schema in use is not exposed to the application."
Also related, Deep Integration of Scripting Languages and Semantic Web Technologies "Instead of loading an ontology, the notion of importing one leads to a different association for the programmer. The difference may be regarded as pedantic or subtle, but it is an important one: the programmer, instead of regarding the ontology as mere data she has to load and access via an API, the imported ontology behaves like a library, extending her possibilities like only code does."
What's covered in the second paper is the problem of the Open World Assumption, in programming they state that we need a third value "unknown" rather than just true and false.
Somewhat related, Hitting reload is the framework job "What's the point in designing tables for a webapp when an RDF-backed store will manage the data for you and RDF queries will come back as tabular data anyway? There are RDF triple stores that will handle in the order 10^6 statements - Leigh Dodds is doing some research on that, up to 10^8 by the looks of things. If I need queries instead of hacking out iterators+fiters I'll use versa/itql/rdql. Now, saying I never want to design another relational schema again is not to say I don't want to use a database. Most of these RDF triple stores are in fact using an RDBMS in the background, as the filesystem and indexer, it's just that the relational schema in use is not exposed to the application."
Sunday, June 19, 2005
Merry Links
* Web Services with WebObjects.
* Why do physicists want to kill their father/grandfather? They didn't prove that you can't sleep with your grandmother though.
* Abraham Bernstein on users "The regular user is not able to cope with strict inheritance."
* x86 OSX is all part of Steve's 10 year master plan.
* MicroSpring "It is a JDK 1.5 IOC implementation in a 30k jar file! It is compatible with the Spring 1.1.3 XML format, but only the IOC elements, and best practice is followed according to the Spring documentation."
* Why do physicists want to kill their father/grandfather? They didn't prove that you can't sleep with your grandmother though.
* Abraham Bernstein on users "The regular user is not able to cope with strict inheritance."
* x86 OSX is all part of Steve's 10 year master plan.
* MicroSpring "It is a JDK 1.5 IOC implementation in a 30k jar file! It is compatible with the Spring 1.1.3 XML format, but only the IOC elements, and best practice is followed according to the Spring documentation."
10 Minute Commits for Better Code
The old way of building software "However an alternative does exist, and it's not that hard to achieve. This alternative involves short development cycles, doing a small amount of tested fully refactored work, and regular commits. There's no magic involved, and it lends itself to fully tested code, constant integration, knowledge transfer and robust code that is easy to update as requirements change (which they always do). Of course if you throw pairing in on top, you get even more of the benefits, however judging by recent experiences this is something that managers and most developers are not ready for, even if it produces excellent results."
In my recent refactoring of JRDF, if it's more complicated that 10 minutes work it doesn't get done. Thinking about how to achieve something in 10 minutes that gets to where you want to go is as powerful a programming tool as anything I've come across. Luckily, running the existing tests in JRDF doesn't take more than a few seconds. Of course, some of this requires a rather powerful IDE too.
In my recent refactoring of JRDF, if it's more complicated that 10 minutes work it doesn't get done. Thinking about how to achieve something in 10 minutes that gets to where you want to go is as powerful a programming tool as anything I've come across. Luckily, running the existing tests in JRDF doesn't take more than a few seconds. Of course, some of this requires a rather powerful IDE too.
Thursday, June 16, 2005
Two Semantic Webs
Bronze from anear, by gold from afar? points to Semantic Web Architecture: Stack or Two Towers? "Features such as closed world assumption and negation as failure (NAF) can be supported by powerful query languages—queries already have a closed world flavour (because distinguished variables can only bind to named individuals), and it is natural to extend this with NAF by way of query subtraction (e.g., the answer to the query “faculty(?x) and NAF professor(?x)” can be computed by subtracting the answer to the query “professor(?x)” from the answer to the query “faculty(?x)”). These features are already supported in query languages such as SPARQL [14] and nRQL [8] (the query language implemented in the Racer system). Moreover, recent work on integrating rules with OWL suggests that future versions of this framework could include, e.g., a decidable subset of SWRL, and a principled integration of OWL and Answer Set Programming [5, 12, 13].
On the other hand, adopting Datalog rules (and DLP with Datalog semantics) would effectively establish two Semantic Webs, with little or no semantic interoperability between the rules based Semantic Web and the ontology based Semantic Web, even at the RDF level. These two versions of the Semantic Web would inevitably be in competition with each other, and this would make the Semantic Web much less appealing: new users would be presented with a difficult choice as to which part to choose, and in choosing would sacrifice semantic interoperability with the other part."
On the other hand, adopting Datalog rules (and DLP with Datalog semantics) would effectively establish two Semantic Webs, with little or no semantic interoperability between the rules based Semantic Web and the ontology based Semantic Web, even at the RDF level. These two versions of the Semantic Web would inevitably be in competition with each other, and this would make the Semantic Web much less appealing: new users would be presented with a difficult choice as to which part to choose, and in choosing would sacrifice semantic interoperability with the other part."
Tuesday, June 14, 2005
JRDF for Learning
In the past I used another project to practically use trendy new things like patterns, XML, Swing, etc. Similarly, I'm going to use JRDF for the same purpose. Kowari is a bit too big for things like going over to Java 1.5, IoC, mocking (real unit tests), lock free alogirthms (including B-Trees) and a few other things that I want to try. I'm not sure it's possible to have a system that doesn't have transactions but it would be interesting to find out. So basically this is just to let people know to expect some changes in JRDF.
Practically, it might mean an RDF/XML based pull parser, persistent JRDF, and more interesting APIs. I'm convinced that developing web services is too expensive and may implode under its own weight - so maybe something based on netKernel or a REST based framework would be a good idea. At the moment I'm just using it to see how much I can get out of IntelliJ.
Practically, it might mean an RDF/XML based pull parser, persistent JRDF, and more interesting APIs. I'm convinced that developing web services is too expensive and may implode under its own weight - so maybe something based on netKernel or a REST based framework would be a good idea. At the moment I'm just using it to see how much I can get out of IntelliJ.
Friday, June 10, 2005
A Reminder About Incremental and Test Driven Development
Iterative and Incremental Development: A Brief History "Project Mercury ran with very short (half-day) iterations that were time boxed. The development team conducted a technical review of all changes, and, interestingly, applied the Extreme Programming practice of test-first development, planning and writing tests before each micro-increment. They also practiced top-down development with stubs."
"We were doing incremental development as early as 1957...where the technique used was, as far as I can tell, indistinguishable from XP...All of us, as far as I can remember, thought waterfalling of a huge project was rather stupid, or at least ignorant of the realities...I think what the waterfall description did for us was make us realize that we were doing something else, something unnamed except for “software development.”"
Other notable references include Boris Beizer and Bill Hetzel in "Introduction to Test Driven Development of Embedded Systems Software": "The value of writing tests early in the design process was first mentioned in 1980 by Boris Beizer, a noted software testing expert. He described the benefit to the testing process of thinking about testing earlier, and then elaborated this point to include the idea that tests developed before the targets of the test may provide additional value in guiding the design effort. Decades later, we have come to believe this is true for a variety of reasons."
"We were doing incremental development as early as 1957...where the technique used was, as far as I can tell, indistinguishable from XP...All of us, as far as I can remember, thought waterfalling of a huge project was rather stupid, or at least ignorant of the realities...I think what the waterfall description did for us was make us realize that we were doing something else, something unnamed except for “software development.”"
Other notable references include Boris Beizer and Bill Hetzel in "Introduction to Test Driven Development of Embedded Systems Software": "The value of writing tests early in the design process was first mentioned in 1980 by Boris Beizer, a noted software testing expert. He described the benefit to the testing process of thinking about testing earlier, and then elaborated this point to include the idea that tests developed before the targets of the test may provide additional value in guiding the design effort. Decades later, we have come to believe this is true for a variety of reasons."
Monday, June 06, 2005
Drools 2.0
Drools 2.0 Released "Drools is designed to allow pluggeable language implementations. Currently rules can be written in Java, Python and Groovy. Drools also enables Domain Specific Languages (DSL) via XML using a Schema defined for your problem domain. DSLs consist of XML elements and attributes that represent the problem domain. An XML Authoring tool provides a semi-rapid development environment with a drag and drop type interface based on the provided Schema."
Examples: House Example and Semantics Module Framework.
Searching around for a screenshot came across SEMANTIC CONFLICT DETECTION IN META-DATA – A RULE BASED APPROACH.
Examples: House Example and Semantics Module Framework.
Searching around for a screenshot came across SEMANTIC CONFLICT DETECTION IN META-DATA – A RULE BASED APPROACH.
Friday, June 03, 2005
Keh-nig-it
* "The SEMANTIC Knight always triumphs! Have at you! Come on, then. I have an battalion of KR theorists on my side". Via Too Close To Home
* Google Sponsored Semantic Web API abstraction project "While a lot of semantic web APIs are available (Jena, Sesame, Kowari etc..) especially for the java language, there is no standard set of interfaces and wrappers so that middle ware RDF toolkits or higher level RDF based API can be built regardless of the underlying api/srdf storage. The project will certainly not start from scratch, but rather from earlier discussions and code to build upon (See jrdf, classes in the Simile projects etc..)."
* Tetris Shelving.
* Google Sponsored Semantic Web API abstraction project "While a lot of semantic web APIs are available (Jena, Sesame, Kowari etc..) especially for the java language, there is no standard set of interfaces and wrappers so that middle ware RDF toolkits or higher level RDF based API can be built regardless of the underlying api/srdf storage. The project will certainly not start from scratch, but rather from earlier discussions and code to build upon (See jrdf, classes in the Simile projects etc..)."
* Tetris Shelving.
Subscribe to:
Posts (Atom)