Tuesday, February 10, 2004

RDF in GPL and LGPL

Creative Commons Includes GPL And LGPL Metadata "I was looking at the Creative Commons site this weekend, and was surprised to find, on their license generation page, entries (translated into Portuguese) in a sidebar for the GNU General Public License and GNU Lesser General Public License, including RDF blocks. Since CC is pushing for projects that can generate, validate, display and search for CC license metadata, how cool would it be to be able to do a Google search for GPL-licensed material, or a P2P network for MP3s released under the CC Attribution-ShareAlike license? As an example, Nathan Yergler has released mozCC, a plugin for Mozilla and Firebird that allows you to view CC license information embedded in a webpage, and provides icons on the status bar displaying the CC license options."

Time to do one for MPL I guess.

RAD XUL

Building RAD Forms and Menus in Mozilla "In Rapid Application Development with Mozilla, Web, XML, and open standards expert Nigel McFarlane explores Mozilla's revolutionary XML User interface Language (XUL) and its library of well over 1,000 pre-built objects. Using clear and concise instruction, McFarlane explains what companies such a AOL, IBM, Hewlett-Packard, and others already know—that Mozilla and XUL are the keys to quickly and easily creating cross-platform, Web-enabled applications. The Mozilla Platform encourages a particular style of software development: rapid application development (RAD). RAD occurs when programmers base their applications-to-be on a powerful development tool that contains much pre-existing functionality. With such a tool, a great deal can be done very quickly. The Mozilla Platform is such a tool."

Thursday, February 05, 2004

SemEarth

The Semantic Earth "Thanks to the constellation of technology that enables digital networks to be laid over the places of the earth, wherever we are we will be able to hear the human conversation that has occurred about that place - the history that occurred there, the aesthetics to be savored, the commerce transpiring at that very moment, recommendations offered by strangers and friends."

More comments: "Predictably, I'm curious because this is an area that my research group has been interested in for some time. For example, our work in creating digital representations for people, places and things led us to semantic location and the Websign project...will a proprietary model be operative?"

Planetarium

XML Watch: Planet Blog "As the number of Planet-style aggregators grows (while I'm writing this, Planets Apache and SuSE are under active development), so grows a variety of software for creating the aggregated sites. There are now at least three codebases for creating such sites, originating with Monologue, Planet GNOME, and Planet RDF. It would be good if each of these codebases could interoperate at least on the basis of configuration files, such as the RDF blog listing from Listing 1. Additionally, we may want a more advanced way of describing each of the planets, perhaps so an über-aggregator -- the Planetarium! -- can be made. (Actually, Jeff Waugh, who created Planet GNOME, has just registered "planetplanet.org", so watch this space!)

I'll leave you with the code in Listing 5, which is a suggestion of how multiple planets could be described; processors follow the seeAlso links to retrieve a list of contributors for each planet. If the choice is made to use RDF/XML, creating the über-configuration file is as easy as aggregating the various RDF blog lists."

A great article and all, especially Figure 2.

Quick Links

* The Semantic Web Made Easy - about a startup using RDF called Radar Networks.
* Protege 2.0
* SemWeb Central - OS Semantic Web tools.
* OntoJava (linked to previously)
* Mac G5 Cude - Instructions on how to build a Mac cube with the G5's grill look.
* Why your Movable Type blog must die - Another anti-blog rant.
* What if there was a data format, and nobody cared? - About social networking and RSS.

Wednesday, February 04, 2004

Docco

Tockit is an open source project which was written by people from DSTC. It has a great interface that uses lattices called Docco. It's a representation of Lucene results combined with metadata extracted by plugins like POI. It seems to be only using the text found in documents, not actual concepts.

It helps to remove the empty nodes to better understand the results.

Screenshots are also available.

Other FCA tools written using Java are available at the ToscanaJ site.

SWOOP

SWOOP (Semantic Web Ontology Overview and Perusal) "SWOOP is a simple and elegant utility to browse Ontologies written in OWL in a hyperlinked thesaurus-style format...Allows users to add OWL ontologies to the Knowledge Base and browse terms listed in them (sorted alphabetically). These ontologies can be saved locally for faster retrieval at a later stage (uses Jena 2.0)"

EII

A new view on data "Composite Software is at the forefront of this trend. Here's what EII is not, according to Jim Green, Composite's CEO and chairman: EII is not EAI (enterprise application integration). EAI pushes data around; EII is a pull system. If an address changes in your CRM application, EAI will push that information out to your ERP (enterprise resource planning) system, for example. Conversely, EII, pulls only the data you need out of ERP and CRM systems and offers it up in a single view for analysis.

Composite Software's Information Server stores the meta data (the data about the data), the fields, and the relationships between them. When the user executes the query, it fetches the data from the underlying systems to present a synthetic view.

"Composite Software's Composite Information Server joins data from different types of resources and creates an alias so it looks different than when it was stored," Green says.

EII is also not BPM (business process management). It has nothing to do with changing business processes. Composite Software's Murthy Nukala, vice president of marketing, pointed out some of the benefits of EII.

"Data takes up most of the cost of an integration effort. Increased understanding of that data is mission-critical, and you need to have a strategy about how you handle data," said Nukala. "

Tuesday, February 03, 2004

MOFman Prophecies

"If done correctly, a meta data-based approach can allow for reuse of interface definitions, messages and other pieces of the integration puzzle. For reuse to happen, though, developers must be able to quickly and easily find previously developed pieces. A centralized repository approach is one way to let that happen...The Holy Grail for this technology is creating, managing and reusing meta data models that are then redeployed to execution engines to reduce the amount of code that needs to be developed manually."

"Vendors, including Redwood City, Calif.-based Informatica Corp. and Bournemouth, U.K.-based Adaptive, have implemented MOF for their warehouse tools and engines. Informatica’s SuperGlue helps to put context around information used in traditional enterprise application integration (EAI)...[it] collects, stores and helps to analyze meta data, including providing audit trails. “The value is in being able to see the dependencies and linkages between data,” Poonen explained. Six customers, including one federal government agency that Poonen could not name, currently use SuperGlue."

"Beyond technology, successful meta data-based integration will depend on corporate culture and practices. Not everyone agrees that a centralized “center of excellence” approach to integration meta data is an absolute given to make it work, but some methodology or approach clearly is -- especially if reuse is ever going to happen."

"By 2005, predict Gartner analysts, more than half of large organizations will have multiple sources of integration technology. “As that proliferation occurs, being able to recognize the use of meta data and have consistent use across all the different deployments becomes important,” said Thompson."

The next step for meta data: Application integration

Monday, February 02, 2004

Lego as an Analogy

We were just talking about it the other day, I recently came across this Lego diagram describing the trade-offs of business objects.

Saturday, January 31, 2004

Pragmatic Metadata

Content-aware searching "Brin’s pragmatic stance sharply opposes the idealistic view of the Web’s inventor, Tim Berners-Lee, who continues to evangelize his vision of a Semantic Web full of carefully encoded content that we can precisely search and fluidly recombine. My own humble contribution to this debate is a prototype search engine, now running on my Weblog, that tries to steer a middle course between the Scylla of simple fulltext search and the Charybdis of unwieldy tagging schemes and brittle ontologies."

"Remember, the pools of HTML content that your people routinely create, and the infinitely vaster pools to which they have access, are full of intrinsic metadata — including the links, tables, images, and other elements that occur naturally within HTML content. Mining that metadata may be more practical than you think."

Web Services then the Semantic Web

The Web within the Web "All these new protocols, SOAP in particular, took years to develop. Indeed, they're still works in progress, in part because contributing companies want to receive patent royalties or just don't want a competitor to control a standard. Those same concerns sabotaged two earlier transport mechanisms, one from the Unix world and one invented by Microsoft. "

"Although Web services allow a machine to publish its data, making it available to another machine, the two have to agree on the structure of the data they are publishing. In the semantic Web, this sort of agreement will be largely unnecessary. "

Bossam

Bossam rule engine v0.7b42 is available for download. Only binary form of the engine is available, and you need a Java runtime (J2SE 1.4 or later) to run the engine. The engine is not feature rich, has many problems, and is buggy. But still, you can process RDF queries, and perform reasoning over OWL ontologies with Bossam. Currently, Bossam supports only one rule language, Buchingae.

Friday, January 30, 2004

Datacentric Web - Microcontent de ja vu?

The Data Centric Web "This shift in focus from documents to data and from humans to computers is simple and yet profound. Just imagine a world in which every piece of data is immediately and automatically accessible from any computer via the web using a simple, universal set of protocols and formats. Indeed, such a vision has long represented the ?holy grail? of Enterprise Application Integration (EAI) and yet attempts to realize this vision have been woefully inadequate to date."

Very similar to an older article about the microcontent client and of course very similar to the Semantic Web too.

Introducing the Microcontent Client
"The microcontent client is an extensible desktop application based around standard Internet protocols that leverages existing web technologies to find, navigate, collect, and author chunks of content for consumption by either the microcontent browser or a standard web browser. The primary advantage of the microcontent client over existing Internet technologies is that it will enable the sharing of meme-sized chunks of information using a consistent set of navigation, user interface, storage, and networking technologies. In short, a better user interface for task-based activities, and a more powerful system for reading, searching, annotating, reviewing, and other information-based activities on the Internet."

The good and the bad

I'm in heaven "Today I'm in geek heaven in so many ways. I read Curtis Hovey's recent weblog entry. He writes:

I strongly feel that a GNOME metadata solution should be based on metadata standards: RDF, OWL, FOAF and use common grammars."

From the original posting: "I strongly feel that a GNOME metadata solution should be based on metadata standards: RDF, OWL, FOAF and use common grammars. I'm shopping for a new Medusa backend because I don't think Medusa should be in the DB business, and it needs an extensible schema."

"The bad

* BDB 4: will everything break when BDB 5 comes out
* Mysql: a bit of a nuisance to setup for single users
* query: applications and users need robust searching
* scalability: will this work at 100 megs, the size of my Medusa db"

Taxonomy Warehouse

Agency taxonomies are a tall order, experts say "An agency building an enterprisewide taxonomy should expect to see more than a million categories within their design, according to Claude Vogel, chief technology officer for search engine company Convera Corp. of Vienna, Va."

"“People underestimate the magnitude of how big their taxonomies will be,” Vogel said, adding that commercial software, such as Convera’s, can handle most, though not all, of the job. "

http://www.taxonomywarehouse.com/

Tuesday, January 27, 2004

Two Interviews

Checking in with the Inventor: Tim Berners-Lee ""The general public is seizing on the Web as a way to have a conversation," he said in our own chat this week. "That for me is very inspiring. It doesn't tell me something about the Web. It tells me something about humanity. The hope for humanity is that people do want to work things out. They do want to come to common understandings, and they will do it by constantly refining the way they've expressed their own ideas--and occasionally, on a good day, listening to the way other people have expressed theirs.""

Under the Iron interview with Aaron Swartz "You’ve put a tremendous amount of work in, for example, RDF and RSS 1.0 (the latter using the former). People say this is the basis of the “Semantic Web”. Could you cue us in on what they hope to achieve with this, how they will make everyone start doing something to achieve it, and what exactly it is we’ll start doing? Do you believe this is possible?

So, uh, here’s the plan:

1. Collect data

2. ???????

3. PROFIT!!!

Uh, more specifically, the idea is to get everyone sharing their vast databases of information in RDF with each other. Then we can write programs that put this data together to answer questions and take actions to make our lives easier."

Thursday, January 22, 2004

What makes technology succeed?

What matters? "In many cases, it is much more important that a choice be made so that we as a society can benefit from the network effects, than it matters which choice is made...the choice between technologies is often of much smaller significance for us as a society than that there be a choice. Networks effects have their role in this and provide many of the benefits.

But if we want to benefit not just from the network effect but also from the advantages of technology, it is in everyone's interest that the network effects cut the right way: that we choose as a society the technologies that work best.

Now, if network effects are the best predictor, then we must infer that the people who actually are responsible for making a good decision are the early adopters. In IT, that means you. You have a responsibility to judge what matters not by network effects but by technical merit. This is a special case of the Categorical Imperative of Immanuel Kant, which you may dimly remember is phrased something like this: “act only on that maxim by which you can at the same time will that it should become a universal law,”1 but which your mother may have expressed more colloquially as, “What would the world be like if everyone did that?”"

"But the reasons that it is a good idea for them to be widely adopted have nothing to do with the differences between SGML and XML, and everything to do with the essential characteristics of the languages...But the choice of any technology is a cost/benefit calculation. And the only changes XML made to that calculation were in lowering the costs of deployment, not in adding any benefits — unless you count the the benefits of the network effect, which are, as I have suggested, considerable."

OWL "Tiny"

Discussion on Owl Tiny.

Wednesday, January 21, 2004

Cluster Graphing

"Anyone who has ever had to complete a what doesn't belong question on a test has an interest in clustering technology. How close are the terms "slime mold", "skunk odor removal" and "luxury bathroom" anyway? Zoom in here and find out. (Clue: They are all green)."

http://labs.yahoo.com/demo/clustergraph/top_level#img

Okay, that looks cool...then put your mouse over the map...

Kowari Update

* iTQL will be changing to be RDQL with proprietary extensions. This is based on the recent RDQL submission and previous discussions we've had.
* The Kowari lite development (just the minimum number of jars to get it going) is continuing. The iTQL command line UI has been improved so it's at least as good as the web iTQL UI.
* RDF Query Languages Eric van der Vlist's question, "Where are the triples?", we've often thought the same thing. The iTQL interpreter has planned, for a long time, to support spitting out RDF/XML results as well as it's existing XML and ResultSet based answers.
* CVS is going to stay internal until we get significant external development. It will be updated infrequently (as bugs are fixed) and with each release.
* Started looking at Aquamarine's API as far as JRDF is concerned. JRDF has had some minor updates (only in CVS at the moment).

Actually, after thinking about querying what you really want is to define a query and have it return all the RDF/XML related to a resource that matches the query - it's something that Guha wrote about (as given by Dan Brickley). This would require OWL and schema support but it's something that is unique over an existing SQL database. You could also, hack this up, by creating a large WHERE clause that defines all the properties of the resource but that's a lot of effort to do each time.

Semantic Google, Semantic Web

Reading this discussion about semantics (linked to by this blog entry) it's especially encouraging to see people when they are talking about this stuff to talk about RDF and the Semantic Web.

Especially, this on the second page:
"Then I re-read G's analysis of Vijay's article (previous post in this thread) in which he points out that Google/Froogle is already extracting this semantic information from non-RDF documents and doing a pretty good job of it all things considered (even if they are trying to sell off a forum moderator, and pretty cheaply, too, I might add ).

So, if I'm understanding things correctly... we don't have to convert everything over to XML (at least not right away) in order for this to work. Which is a good thing, because there are a buncha individuals and mom & pops out there (and some companies who ought to know better and could certainly afford the upgrade) who haven't even started using CSS and HTML4.x, much less XHTML or XML."

"Think of RDF (Resource Description Framework) as the foundation. Think of XML (eXtensible Markup Language) as the formatting language used to deliver it."

Putting Ontologies to Work

Judging the likely Success of an Ontology "Clay Shirky is obviously right when he states that a single monolithic ontology will never work. His critics are equally right when they claim the Semantic web will only work if it is a m�lange [melange] of multiple interoperable Ontologies. What is missing from the debate is a more detailed explanation of what ontologies are good at, how they interoperate, and why systems based on ontologies succeed or fail."

"Ontologies, far from being an unproven new concept, are already in practical daily use. They form the foundation of classification systems, databases, and object oriented software applications."

I enjoyed reading this article, especially as it touches on many areas (like the relational model) and previous articles. This is virtually the perfect anti-Shirky piece.

Tuesday, January 20, 2004

Apache XML

Not what you would initially think: ""Within the Apache Longbow, eXtremeDB will manage secure, digitized battlefield data. eXtremeDB's XML interface will facilitate communications, both internally and between the attack helicopter and external (ground and air) systems. Embedded software including eXtremeDB will run on airborne PowerPC processors and a commercial real-time operating system (RTOS). The program is being developed by The Boeing Company's Phantom Works organization in Mesa, Ariz."

http://www.xmlmania.com/news_article_846.php

Combine Two Technologies

Like putting a clock in an existing product I keep thinking of combining an XML Swing library with JSF. It would be nice to use JSF as a way to provide the abstraction for differing rendering technologies (this was the first thing I thought when I saw JSF). This interview had some interesting bits of information on this:
"One of the unique things about Faces is that it allows you to have separate classes for rendering a UI component. So a simple text box can consist of a UIInput component, which represents the concept of collecting user input, and a Text renderer, which knows how to display a textbox in HTML. You can create separate renderers for different types of clients -- one for HTML, one for SVG, and one for WML, for example...the third-party component market will continue to grow, not just with HTML components (which will be first), but also components and renderers that support other devices and richer clients."

"There's a sample in the current Faces early access release of XUL instead of JSP, but I think more work needs to be done to prove that other display technologies can really be first-class citizens."

"JavaServer Faces is also a good technology for thin client applications that aren't HTML-based. I've mentioned WML, but you could also write a Java applet application, or some other non-browser client that works with JSF. We'll see these types of applications evolve over time.

Personally, I think fat clients are great for some applications, like RSS News Readers. But web applications are great for other things, and JSF is a good way to build those types of applications. "

The latest download includes XUL in the non-JSP examples.

Monday, January 19, 2004

Related To

Semantic Similarity in a Taxonomy "This article presents a measure of semantic similarity in an IS-A taxonomy based on the notion of shared information content. Experimental evaluation against a benchmark set of human similarity judgments demonstrates that the measure performs better than the traditional edge-counting approach."

Xen

'Xen' programming language unites C#, XML and SQL programming languages. ""I am currently working on language and type-system support for bridging the worlds of object-oriented (CLR), relational (SQL), and hierarchical (XML) data, and of course first class functions," explains Mejer."

Mejer's paper, Unifying Tables, Objects, and Documents explains this idea more fully. Beware the Haskell programmer. Some of the syntax reminded me of Groovy. ExtremeTech also have an article.

Which Schemas?

WinFS Is a Storage Platform "WinFS is an active storage platform for organizing, searching for, and sharing all kinds of information. This platform defines a rich data model that allows you to use and define rich data types that the storage platform can use. WinFS contains numerous schemas that describe real entities such as Images, Documents, People, Places, Events, Tasks, and Messages."

Saturday, January 17, 2004

No unstructured data

MORE ON “UNSTRUCTURED” THINKING "There is no such thing as "unstructured data". That means random noise, which has no structure whatsoever and, therefore, is meaningless. It is the structure that gives meaning/content and makes data.

It has nothing to do with scanning, or incompleteness, or missing, or anything. It is structured, whatever it is. Diagrams have one type of structure, partial documents different types of structure, but there is always some structure by definition.

The term "unstructured data" is a misnomer based on misconception: it essentially refers to data that is not structured in tables, or spreadsheets, or whatever; mainly text, graphics, etc. But that is not unstructured, it's just different structures than tables or spreadsheets, that's all.

And that's a core issue, because structure determines the integrity and manipulation of the data, which are different for each type of structure. The point of relational structure is that it is the simplest formal structure for integrity and manipulation. Any other structure adds complexity, but no power."

Thursday, January 15, 2004

Another Java RSS Parser

FeedParser "The main API is very similar to JAXP, TRaX, SAX, and is designed to be very flexible. Having been a veteran of the RSS wars, member of the RSS 1.0 working group, and Atom developer, I think this takes into consideration all major issues with RSS/Atom feed formats and integration with the Java language...RSS serialization support. Serializers for all RSS versions (1.0, 2.0, Atom, etc) with the same code."

Kowari

Well, it's not quite there yet but it will be available at SF's Kowari Project page (a 45MB download and in CVS). I think I've mentioned this before, one of the future goals is to reduce the download to the minimum set of jars.

Tuesday, January 13, 2004

Why the Dock Sucks

Top Nine Reasons the Apple Dock Still Sucks "The Dock is like a brightly-colored set of children's blocks, ideal for your first words—dog, cat, run, Spot, run—but not too useful for displaying the contents of War and Peace."

Active Internet

A collection of articles about how computers and the Internet are turning people into content producers not just consumers:
* Weblogs, RSS and the Rise of the Active Web - "...we show how blogging – originally a cross between self-expression and journalism – and its tools have morphed to give users some of the power promised by the so-called Semantic Web...they can construct personal news or commerce portals for themselves or for third parties, track multi-person blog conversations across the Web, or figure out other ways to control their digital environment that we have not thought of yet."
* The New Economy Hack: Turning Consumers into Producers - "That industry lately has become vigilant about threats from its customers, which it still thinks of as consumers. Instead it should be watching how Apple transforms those consumers into producers."
* Democratizing the Media, and More - "Smarter folks will understand the enormous opportunity it represents. They can start listening, really listening, to what people are saying. And they can dip into the vast pool of creative talent that exists outside the usual channels."

I finally have an excuse to link to Bush In 30 Seconds. My favourites were: In My Country, What are we teaching our children?, Imagine, Human Cost of War and Bush's Repair Shop. I still think the quality, even in the top 14, was spotty but that's where you need good annotation and recommendation software.

Monday, January 12, 2004

The Importance of Ontology

Ontology and Integration - Managing Application Semantics Using Ontologies and Supporting W3C Standards " Ontologies are important to application integration solutions because they provide a shared and common understanding of data (and, in some cases, services and processes) that exists within an application integration problem domain, and how to facilitate communication between people and information systems. By leveraging this concept we can organize and share enterprise information, as well as manage content and knowledge, which allows better interoperability and integration of inter- and intra-company information systems. We can also layer common ontologies within verticals, or domains with repeatable patterns."

Sunday, January 11, 2004

The old fashion way of integration

Compare and Contrast JOLAP and XML for Analysis and Intelligent Business Strategies: OLAP in the Database. I've covered some of this previously. Especially, JMI.

"JOLAP is a J2EE objected-oriented application programming interface (API) designed specifically to addresses the programming needs of Java developers by providing a standard set of object classes and methods for BI."

"XMLA (www.xmla.org) is a linguistic interface with no preference for programming language or object model. This linguistic interface is implemented as a web services and also defines a standard query language (mdXML) for BI."

"Hyperion views JOLAP and XMLA as complementary rather than competing standards. Although you can implement XMLA without using JOLAP, the JOLAP specification supports the web services architecture that depends on J2EE application servers, XML, and SOAP messages.

In fact, Hyperion’s implementation of XMLA uses our Java API (which was developed based on our JOLAP specification work) to communicate with the Essbase Analytic Services (OLAP Server). Our XMLA web service accepts a SOAP message, takes the mdXML statement contained in the SOAP message, and passes it to the Analytic Services engine for processing through the Java API. The result set is passed back to the XMLA web service through the Java API, where it is wrapped in a SOAP message and sent to the requesting client."

Friday, January 09, 2004

Blogs are bad, don't do blogs

Why I Fucking Hate Weblogs! "Weblogs suck ass. What the fuck is up with this shit? Fuck. Who the fuck cares what these people think about oatmeal or what the UN did last week? Nobody! Who reads these weblogs? Nobody! Maybe fellow weblog authors read each others weblogs out of a sense of desperation...the feeling that if they read someone else's weblog, someone will read theirs. It's kindof like cooperative advertising too, people will cross-post, linking weblog entries to each other's weblogs. How fucking pathetic is that? I hate weblogs. "

It's convinced me...this has been a waste of time.

Perfect Company

"And if you did create the perfect organisation – perfect in organisational terms, that is, one that would magically hoover up all of the money and destroy its competitors, as soon as you achieve that perfection you would also achieve destruction – because our society seems to thrive best where there are many ideas contending and where no one organisation/form of government/set of ideas has eliminated all the rest."

Interview with Martha Atwood

Polite Society

What you can't say "It seems to be a constant throughout history: In every period, people believed things that were just ridiculous, and believed them so strongly that you would have gotten in terrible trouble for saying otherwise.

Is our time any different? To anyone who has read any amount of history, the answer is almost certainly no. It would be a remarkable coincidence if ours were the first era to get everything just right."

What you can say.

Thursday, January 08, 2004

Relational Web Services

XQuery on the Web talks about Xquery: Meet the Web.

Dare quotes: "In fact, this separation of the private and more general query mechanism from the public facing constrained operations is the essence of the movement we made years ago to 3 tier architectures. SQL didn't allow us to constrain the queries (subset of the data model, subset of the data, authorization) so we had to create another tier to do this.

What would it take to bring the generic functionality of the first tier (database) into the 2nd tier, let's call this "WebXQuery" for now. Or will XQuery be hidden behind Web and WSDL endpoints?"

And responds with:
"Every way I try to interpret this it seems like a step back to me. It seems like in general the software industry decided that exposing your database & query language directly to client applications was the wrong way to build software and 2-tier client-server architectures giving way to N-tier architectures was an indication of this trend."

"Data model subsets" - don't you mean views?

Also Dare says, "All this indirection with WSDL files and SOAP headers yet functionality such as what Yahoo has done with their Yahoo! News Search RSS feeds isn't straightforward. I agree that WSDL annotations would do the trick but then you have to deal with the fact that WSDL's themselves are not discoverable."

Which I would refer anyone interested to the paper in the JOWS: Automated Discovery, Interaction and Composition of Semantic Web Services.

XML For You and Me, Your Mama and Your Cousin Too "At this point if you are like me you might suspect that defining that the web service endpoints return the results of performing canned queries which can then be post processed by the client may be more practical then expecting to be able to ship arbitrary SQL/XML, XQuery or XPath queries to web service end points.

The main problem with what I've described is that it takes a lot of effort. Coming up with standardized schema(s) and distributed computing architecture for a particular industry then driving adoption is hard even when there's lots of cooperation let alone in highly competitive markets."

Journal on Web Semantics

Journal on Web Semantics "The Journal on Web Semantics and also this Website is approaching scientific publishing from a different angle: our topic demands more than just the production and printing of papers, but also the distribution of ontologies and running code. An early slogan of W3C standardisation efforts was 'rough consensus and running code' - this applies also for the Semantic Web - maybe changed to 'rough consensus, running code, and ontologies'."

Wednesday, January 07, 2004

Followup on Relational RSS

RSS, old enough to be having relations? "Seb mentions several of the operators of Codd's relational algebra, and, it seems to me there are two general reasons why everyone isn't already operating on RSS as relational data: 1) it is distributed across many files, and 2) the hierarchic XML structue of RSS."

"The main issue I am dealing with now is what types of data structures and formats work best with the various combinations of uses between data interchange, data storage, and querying."

Okay, now I'm convinced that this really is replicating RDF and I would have to encourage anyone considering this to pick up an open source RDF library (like Jena or Redland) and use it to perform these operations on RDF based RSS.

The problems that are highlighted are the same ones that various implementations of RDF have had to solve. A problem with serializing a graph (relational data) in XML - that's RDF/XML and it's use of striping. Distributed across many files and being able to search it - that's usually a problem for RDF data stores (like Kowari or other freely available ones).

For example, to get all the documents (blog entries, etc.) authored by Sam Ruby (this is from a previous post describing iTQL): "select $creator subquery( select $type from <rss_schemas> where $type <http://www.w3.org/2002/07/owl#sameAs> <http://purl.org/dc/elements/1.1/creator> ) from <rss_feeds> where $creator $type 'rubys@intertwingly.net';"

Where "<rss_feeds>" can be any number of URIs combined with logical operators.

Of course, you'll need to convert some feeds from XML to RDF. While I often link to RDFT, the more usual ways include XSLT and programmatically using an RDF API and an XML library. One of the quickest ways I've found is using a combination of Jena and Jakarta Apache's XML Commons Digester.

Also related, Base data: relational, RDF, XML.

George W. vs Hitler

George Bush & Adolf Hitler "The internet is littered with pictures of George Bush with a swastika on his chest, a drawn-on mustache, and his arm raised in the Nazi salute. It is therefore no surprise that someone would choose this theme for their video. But why? Has George Bush done anything to justify the comparison? Consider these points:

* Hitler slaughtered six million. George Bush only killed nine thousand or so in Iraq and many fewer in Afghanistan. Hardly a fair comparison.
* Hitler rounded up and killed homosexuals. George Bush only denies them the right to marry. Again, no comparison.
* Hitler rounded up and killed those with physical or mental infirmities. George Bush only cut their medical benefits. No comparison.
* Hitler invaded his neighbors and overthrew their governments. George Bush only invaded and overthrew the governments of two countries, and they were not neighbors. No comparison.
* Adolph Hitler believed in a "master race." George Bush believes in a master religion. No comparison.
* Adolph Hitler was an eloquent and persuasive madman. George Bush is neither eloquent nor persuasive. No comparison
* Adolph Hitler's government was in tight control of its citizens. George Bush has only limited our right to privacy, free speech, and access to lawyers.
* Adolph Hitler demonized Jews. George Bush only demonized Osama bin Laden and Saddam Hussein (with the leaders of Syria, Iran, and North Korea held in abeyance). No comparison.

Obviously, George Bush is no Adolf Hitler. To make sure those videos are never seen, we suggest that he confiscate them and start arresting people. This kind of outrage should not be allowed in a free society."

Tuesday, January 06, 2004

Set Theory with RSS

The Algebra of RSS Feeds ""Taking a cue from the operations of set theory," Paquet writes, "we could for instance define the following:

1. Splicing (union): I want feed C to be the result of merging feeds A and B.
2. Intersecting: Given primary feeds A and B, I want feed C to consist of all items that appear in both primary feeds.
3. Subtracting (difference): I want to remove from feed A all of the items that also appear in feed B. Put the result in feed C.
4. Splitting (subset selection): I want to split feed D into feeds D1 and D2, according to some binary selection criterion on items.""

The original post.

While I wouldn't say that RDF has the monopoly on set theory (and operations like union and intersection) it does seem like reinventing the wheel.

Monday, January 05, 2004

Misinformation

Internet creator Berners-Lee knighted "British physicist Tim Berners-Lee, who invented the World Wide Web -- or at least better access to it -- has been awarded a knighthood in London.

Without his creation, there would be no computer addresses, no e-mail and the Internet might still be the exclusive domain of a handful of computer experts, the Independent reported."

RSSOwl

"RSSOwl is a free RSS (0.91, 0.92, 1.0, 2.0) newsreader written in Java programming language using SWT as fast graphic library. Features of RSSOwl include reading RSS or RDF newsfeeds in a comfortable tab folder, save newsfeeds in categories, export them to PDF / HTML or OPML, and view news in an internal browser."

Even though it had some issues with freezing, laying out some of the UI, and other bugs it's easy to get running and it's not too bad. It'd be nice to have support for RSS autodiscovery. I think I still prefer NetNewsWire or NewsMonster although I should try the others.

Sunday, January 04, 2004

More Commercial RSS

* RSSAds "The ad engine for RSS feeds."
* k-collector "k-collector is an enterprise news aggregator that leverages the power of shared topics to present new ways of finding and combining the real knowledge in your organisation...The k-collector archicture combines clients for leading weblogging software with a server based aggregator and web application."

Waypath

"Waypath makes use of Think Tank 23's unique information retrieval platform, Nav4, which automatically analyzes content, such as weblogs, and links documents that share common topics. Using Nav4, Waypath provides both keyword search and contextual navigation of individual weblog posts."

SDK feature list includes things such as: concept-driven similarity browsing, under load, handling over 600 queries/minute and 1 million documents on a typical installation and there's no taxonomy to maintain.

Here are Morenews's related links (top link is Themes and metaphors in the semantic web discussion) and related books.

OS JavaServer Faces

Smile Now includes its own renderkit. Next release will be fully compliant, apparently.

More associated links: JavaServer Faces home page, more details in chapters 21 and 22 of the Web Services Tutorial or alternativately a tutorial for the impatient.

Saturday, January 03, 2004

iBox

Exclusive Insider Information: Apple iBox in production. "The iBox plugs into your TV and acts as a hub for your digital devices and computers. Unlike the EyeTV from Elgato, the iBox is a standalone machine, not something to plug into an existing computer. The iBox can be scheduled to record TV, but unlike TiVos it does not serve as a "what's on and when" service rather a hard drive / media based recording device (new aged VCR). With its built in 802.11b & 802.11g from its AirPort Extreme card, one can access the home folders of any user on any wirelessly networked Mac or PC. The iBox has its own version of the popular iPhoto and iTunes software which is a welcoming plus to Mac OS 10 veterans and easy for Windows users to adopt as well."

Time to sell my Shuttle. A picture you can eat with a spoon. Tasty Apple rumours.

NoodleTools

Finding what you want on the web "Debbie is a trained librarian, and it shows. She understands that a single search engine is never going to do everything, no matter how good its indexing or how large its database."

"I do not think we will ever solve the search problem until we move away from the dumb web we have today towards something like the semantic web, a project that Sir Tim Berners-Lee has been pushing ever since the first web conference in 1994.

Once links carry meaning then it will become more of a distributed database than the vast heap of unstructured documents we have built so far.

And once we have a database then we can classify, index and search it properly."

NoodleTools.

Analysis Engine

A Fountain of Knowledge "...imagine a marketing researcher trying to find out the online attitude of consumers toward the popular rock singer Pink. The researcher would have to wade through an ocean of search results to sort out which Web pages were talking about Pink, the person, rather than pink, the color.

What such a researcher needs is not another search engine, but something beyond that—an analysis engine that can sniff out its own clues about a document’s meaning and then provide insight into what the search results mean in aggregate. And that’s just what IBM is about to deliver. In a few months, in partnership with Factiva, a New York City online news company, it will launch the first commercial test of WebFountain...Up to now this kind of aggregate analysis was possible only with so-called structured data, which is organized in such a way as to make its meaning clear. Originally, this required the data to be in some sort of rigidly organized database; if a field in a database is labeled “product color,” there is little chance that an entry reading “pink” refers to a musician."

"Although the pooled data is compressed to about one-third its original size to reduce storage demands, WebFountain still requires a whopping 160 terabytes plus of disk space. It uses a cluster of thirty 2.4-GHz Intel Xeon dual-processor computers running Linux to crawl as much of the general Web as it can find at least once a week."

"WebFountain’s builders admit it’s not always able to guess right, but they point out that humans can also be confused by ambiguous meanings."

"Because the data has been converted from an unstructured format to a structured XML-based format, IBM and its partners can fall back on the data-mining experience and methodologies already developed for analyzing databases. The structured format also provides an easy target for developing new analytic tools."

"This, perhaps more than anything else, is why WebFountain looks like a winner. By creating an open commercial platform for content providers and data miners, it will foster rapid innovation and commercialization in the realm of machine understanding, currently dominated by isolated research projects."

Friday, January 02, 2004

Reclaim the Semantic Web

Fight back "For the technologists among us I would recommend you read this piece by Mark Nottingham on reclaiming the Semantic Web from military purposes. We need to stop wasting time on bullshit like Freindster and start using this technology to do something useful like faciltating citizen oversight of the government."

The Semantic Web’s Dirty Little Secret

IBM Emerging Technologies Toolkit

1.2 Released "Version 1.2 contains Service Data Objects (SDO), Policy-Based IT Management Demo, Semantic Web Services, Autonomic Computing Toolset, WS-Manageability demo, WS-Trust, WS-Addressing, Web Services Failure Recovery, and Service Domain technology."

Serendipity Server

One of Danny Ayers' New Year resolutions:
"...text search, creation of triples using machine learning techniques...A server-side tools that combines Semantic Web and machine learning technologies to autodiscover connections between ideas."

It seems similar to some of Tim Bray's search vision Basic Resource Finder:
"BRF will have built in most of the lore on result ranking I wrote up earlier in this series, with the possible exception of Latent Semantic Indexing. Crucially, it will have some facilities to make it easy to feed back popularity and usage counts into the ranking heuristics."

eventSherpa, SW Killer App

eventSherpa - an RDF desktop application for Windows (at last!) "The desktop app is a good looking, user friendly and very functional calendar. Where it starts to get good is that I can publish my local calendars to the eventSherpa Calendar server. They will then be available as HTML and more importantly as RDF feeds using the RDF Calendar schema for anybody to subscribe to either using the eventSherpa client or any software that can consume this RDF vocab."

Thursday, January 01, 2004

OWL Implementations

OWL Implementations (commercial implementations are Cerebra and Snobase) also includes OWL Test Results.

Wednesday, December 31, 2003

Statement vs Stating

This comes up from time to time (at work and on mailing lists) a useful summary of an RDF statement and a stating:
Statements/Statings. From the ILRT Semantic Web technical reports at a glance. Two other references: Does the model allow different statements with the same subject/predicate/object? and also part of Reification in the RDF Semantics document.

Commercial Ontologies

TeraView, Level 5 and BioWisdom starting 2004 in style "Ontology specialist BioWisdom also plans to make a “big announcement early in the New Year.

Ontology is a branch of science that deals with knowledge capture and representation. BioWisdom’s approach involves the development of specialised knowledgebases and the software tools for managing them.

Chief executive Gordon Smith Baxter said: “2003 has been a great year for BioWisdom. In January we secured a £2.5m investment from MB Venture Capital II and Merlin Ventures Fund IV."

BioWisdom and Network Inference on using ontologies for drug discovery.

Friday, December 26, 2003

Some Holiday Links

* Improved Topicalla Screenshot
* Weedshare
* XML 2003 Conference Diary - Notes the continual rise in interest in the Semantic Web.
* SnipSnap 0.5 - Now with the snips available as RDF.

Friday, December 19, 2003

Quintuples

Trust, Context and Justification While I'm not sure about using 5 tuples, we use 4 and make statements about the 4th tuple in order to do things like security, it's still an interesting paper with some good references.

Google Searching for Relevance

A Quantum Theory of Internet Value "When "the Internet" was unveiled to a doughnut-eating public a decade ago, we were promised unlimited access to vistas of encyclopedic knowledge. Every body would be connected to every thing, and we would never be short of an answer. What with the abundance of information, and the costs of transporting information approaching zero, the world would never be the same again.

Of course, a decade on, we know that real economics have prevailed. Information costs money. Those transport costs certainly aren't zero. And faced with a choice of a million experts, people gravitate towards experts with a good track record: i.e., for better or worse, paid journalists, qualified doctors or other centers of expertise.

Taxonomies also have been proved to have value: archivists can justify a smirk as manual directory projects dmoz floundered - true archivists have a far better sense of meta-data than any computerized system can conjure. If you're in doubt, befriend a librarian, and from the resulting dialog, you'll learn to start asking good questions. Your results, we strongly suspect, will be much more fruitful than any iterative Google searches. "

"At a convivial dinner recently, John Perry Barlow asked me why no one had written a story about how the most powerful organisations in the world were dependent on the most awful, antiquated and dysfunctional technology. Well, I ventured (to a deafening silence), maybe they were making ruthless choices, and really weren't too slavish about following techno-fads. Maybe the answer is in the question."

Wednesday, December 17, 2003

Commerical RSS

How to make RSS commercially viable "Without full content no aggregator can add much value by categorizing and filtering infomation, so no purely RSS based aggregator can make much money.

Despite all of the interest around web based syndication, people like Lexis Nexis will still make all the money unless this problem is solved."

Does it? Will it? Must it?

Interview: David Weinberger "What Shelley calls "the semantic web" is the Web itself. She puts it beautifully. And I agree 100% that the Web consists of meaning; it has to because we created this new world for ourselves out of language and music and other signifiers. But that meaning is as hard to systematize and capture as is the meaning of the offline world and for precisely the same reasons. The Semantic Web, it seems to me, often underplays not only the difficulty of systematizing human meaning (= the world) but also ignores the price we pay for doing so: making metadata explicit often is an act of aggression. Human meaning is only possible because of its gnarly, tangly, implicit, unlit, messy context. That's the real reason the Semantic Web can't scale, IMO.

If by "The Semantic Web" you merely mean "A set of domain-specific taxonomies some of which can be knit together to provide a greater degree of automation and improved searching," then I've got no problem with it. It's the more ambitious plans -- and the use of the definite article in its name -- that ticks me off when it comes to The Semantic Web."

Exceptions (again)

13 Exceptional Exception Handling Techniques notes "Declare Unchecked Exceptions in the Throws Clause" and "Soften Checked Exceptions" (always use RuntimeExceptions). This lead to JDO and its JDO Transaction class that uses runtime exceptions (although it does document them) instead of JDBC's use of checked exceptions. Similarly, the Spring Framework and in Chapter 4 of Expert One-on-One J2EE Design and Development the author discusses the usual reasons given to avoid checked exceptions:
"Checked exceptions are much superior to error return codes...However, I don't recommend using checked exceptions unless callers are likely to be able to handle them. In particular, checked exceptions shouldn't be used to indicate that something went horribly wrong, which the caller can't be expected to handle...Use an unchecked exception if the exception is fatal."

With both JDO and Spring the contract offered by the framework tells the client what they can and cannot handle. In my experience, this is not an either or situation. For example in JDO they use "CanRetryException" and "FatalException" - an exception that can be retried, could actually be fatal depending on the context and vice-versa. This often occurs when large frameworks are used in conjunction with one another - at the system integration level. Preventing the developer the choice, when integrating into larger frameworks, what exceptions can and cannot be caught often leads to unexpected exceptions tunneling through layers.

Tuesday, December 16, 2003

Drools in Groovy

Drools (an augmented implementation of Forgy's Rete algorithm) is now available in Groovy.

RDF Matures

Resource Description Framework (RDF) Is a W3C Proposed Recommendation and OWL Web Ontology Language Is a W3C Proposed Recommendation, the next step is Recommendation.

More On Practical RDF

Practical RDF Town Hall "Next, xmlhack editor Edd Dumbill explored how he applies RDF to his personal data integration problems, running personal information through the Friend-of-a-Friend (FOAF) RDF vocabulary, using the Redland framework as a foundation for processing...which has sprouted context features and Python bindings to support this work. "

"In the last presentation, Norm Walsh explained how he was using RDF to make better use of information he already had. Walsh explained that he had lots of data in various devices about a lot of people and projects, but no means of integrating it. Thanks to various RDF toolkits - "just by dumping it into RDF, it just kind of happens for free." Aggregation and inference are easy - and Walsh can get convenient notifications of people's birthdays without duplicating information between a file on a person and a calendar entry noting that."

Monday, December 15, 2003

Corporate Taxonomies

Verity provides standard ways to categorise content "Traditionally, taxonomies have been time-consuming and expensive to set up. A Taxonomy needs to be unambiguous and cover all topics of interest to the organisation. In other words, it has to be Collectively Exhaustive and Mutually Exclusive. Few individuals, not even the company librarian, have the breadth of knowledge of the organisation and its information assets to construct a set of categories that encompasses all information and meets all needs."

"Because a taxonomy reflects the most important knowledge categories of an organisation, organisations that carry out the same business activities need similar taxonomies. (In the same way that such organisations share similar core business processes). This fact and the rising importance of taxonomies to organisations has led Verity to make six tailorable taxonomies available to jump-start the development of an organisation's taxonomy. Verity's six taxonomies suit a range of business activities covering Pharmaceuticals, Defence, Homeland Security, Human Resources, Sales and Marketing, and Information Technology. Organisations that start with these predefined taxonomies can then tailor them to their specific needs. "

Sunday, December 14, 2003

The Winner Takes It All

Power Laws, Discourse, and Democracy "Well, inevitable inequality is one way to characterize the effects of power laws in social networks. But is it the most useful way? Drawing on the same body of research on power laws in social networks, and using similar methods, Jakob Nielsen chose to emphasize instead that, as he put it in a piece published on AlertBox (03.06.16): Diversity is Power for Specialized Sites:"

"Winner-takes-all networks may follow Pareto's Law (the 80/20 rule) with regard to the cumulative distribution of links. But, according to Barabasi in Linked, the distinctive distribution hierarchy of scale free networks will have been broken. Instead, the network takes on what Barabasi describes as a "star topology," in which a single hub snarfs nearly all the links, dwarfing its competitors. "

"It's the the dynamics of emergent systems being formalized in open source. It's the fragile and turbulent architecture of democracy.

By contrast, winner-takes-all networks wipe out the middle ground connecting leaders to the network's other players. With this, winner-takes-all networks strip away the architecture that supports the productivity of local niches."

Saturday, December 13, 2003

More Practical RDF

Practical RDF "There are two features of RDF that I find particularly practical: Aggregation [and] Inference".

Not Influential or Famous

Myths Open Source Developers Tell Ourselves

Friday, December 12, 2003

Groovy is Out

"Groovy is a powerful new high level dynamic language for the JVM combining lots of great features from languages like Python, Ruby and Smalltalk and making them available to the Java developers using a Java-like syntax."

GPath "When working with deeply nested object hierarchies or data structures, a path expression language like XPath or Jexl absolutely rocks."

The SQL and Markup example also looks interesting.

New Java Tools

Algernon-J is a rule-based reasoning engine written in Java. It allows forward and backward chaining across Protege knowledge bases. In addition to traversing the KB, rules can call Java functions and LISP functions (from an embedded LISP interpreter).

JRDF "A project designed to create a standard mapping of RDF to Java."

Google 2005

Searching With Invisible Tabs "Doesn't the future of search look great? Whatever type of information you're after, Google and other major search engines will have a tab for it!"

Highlights that people can suffer from "tab blindness" and why one UI doesn't suite all (fairly obvious).

Greed is Good for Data Emergence

The Age of Reason: The Perfect Knowing Machine Meets the Reality of Content "In brief, the concept of "data emergence" that is central to this knowledge Nirvana is best summed up by James Snell as "the incidental creation of personal information through the selfish pursuit of individual goals." From Snell's perspective, content value is shackled by dumb Web browsers that are used to share information about individuals with Web sites that then try to "personalize" their content - an experience that must be repeated at each and every Web site visited, since this knowledge about individual interests and preferences is not shared site-to-site. Instead of this, the perfect world would have a "smart" content service, probably on one's PC, that would retain knowledge of all of one's personal profile and interests in accessing content; content providers would then be "dumb" sources pumping information into the smart service, not having any detailed knowledge of who is using their services and how. No more nasty Web site publishers, just one perfect tacit machine that knows exactly what you're thinking and allows you to obtain and share thoughts with others."

"Aggregation can happen anywhere to the satisfaction of many."

Kowari Already Out There

Kowari for RDF developers - Early Release It's already out there - found this when doing a Google on Kowari. The real site will be Kowari.org but with OS the source is the real thing I guess.

RSS for the Knowledge Worker

From the Metaweb to the Semantic Web: A Roadmap "At Radar Networks we have been working to define this ontology -- which we call "The Infoworker Ontology" -- with a goal of evententually contributing it to a standards body in the future. The Infoworker Ontology is a mid-level horizontal ontology that defines the semantics of common entities and relationships in the domain of knowledge work -- things like documents, events, projects, tasks, people, groups, etc. The development and adoption of an open, extensible, and widely-used Infoworker ontology is a necessary step towards making the Semantic Web useful to ordinary mortals (as opposed to academic researchers).

By connecting microcontent objects to the Infoworker Ontology a new generation of semantic-microcontent (what we call "metacontent") is enabled. With the right tools even non-technical consumers will be able to author and use metacontent. "

Thursday, December 11, 2003

The Early Days...of the Semantic Web

Early Days Of a Data-Sharing Revolution " And next week, a Chicago company plans to start selling a $36 mini-scanner dubbed "iPilot" that shoppers can use to scan bar codes on products in stores, then upload the data to a computer and compare prices at Amazon.com.

All are examples of how Web sites, relying on a new generation of Internet software, are licensing their databases to business partners and outside developers in an attempt to spark innovation and reach more customers.

"In the past six to nine months, we have started ramping up the program to license eBay's data," eBay Vice President Randy Ching said."

iPilot hey? You'd think with millions of venture capital the least you could do would be better than combining Apple's and Palm's product names.

Saturday, December 06, 2003

Metaweb

The Birth of "The Metaweb" -- The Next Big Thing -- What We are All Really Building "But RSS is just the first step in the evolution of the Metaweb. The next step will be the Semantic Web...The Semantic Web transforms data and metadata from "dumb data" to "smart data." When I say "smart data" I mean data that carries increased amounts of information about its own meaning, structure, purpose, context, policies, etc. The data is "smart" because the knowledge about the data moves with the data, instead of being locked in an application...The Semantic Web is already evolving naturally from the emerging confluence of Blogs, Wikis, RSS feeds, RDF tools, ontology languages such as OWL, rich ontologies, inferencing engines, triplestores, and a growing range of new tools and services for working with metadata. But the key is that we don't have to wait for the Semantic Web for metadata to be useful. The Metaweb is already happening."

Friday, December 05, 2003

More Visualization of RDF

Meta-Model Management based on RDFs Revision Reflection Breaking up the different type of RDF (schema and properties) is interesting although probably still has scaling problems (as with most of these types).

Styling RDF Graphs with GSS "One such solution is GSS (Graph Style Sheets), an RDF vocabulary for describing rule-based style sheets used to modify the visual representation of RDF models represented as node-link diagrams."

I would imagine that somehow taking historgram data and mapping that from graphs maybe more interesting and would scale better. Much like how some image search engines work.

Al Gore - How I would've done it different

FREEDOM AND SECURITY
""I want to challenge the Bush Administration’s implicit assumption that we have to give up many of our traditional freedoms in order to be safe from terrorists...In both cases they have recklessly put our country in grave and unnecessary danger, while avoiding and neglecting obvious and much more important challenges that would actually help to protect the country...In both cases, they have used unprecedented secrecy and deception in order to avoid accountability to the Congress, the Courts, the press and the people." "

"In other words, the mass collecting of personal data on hundreds of millions of people actually makes it more difficult to protect the nation against terrorists, so they ought to cut most of it out.""

Thursday, December 04, 2003

Semantic Merging

Skip This Rant and Read Shirky "Shirky sums up many metadata challenges with a concise statement: "it's easy to get broad agreement in a narrow group of users, or vice-versa, but not both." Hey, if you don't make your metadata structurally interoperable, you can't have semantic merging."

I found the diagram, "Content Enterprise Metadata: Structural Interoperability & Semantic Merging" to be quite instructive.

Wednesday, December 03, 2003

Zeitgeist Mining

On RSS, Blogs, and Search "I've been thinking lately about the role of blogs and RSS in search, and that, of course, has led me to both the Semantic Web and to Technorati, Feedster, and many others. Along those lines, I recently finished a column for 2.0 on blogs and business information. I can't reveal my conclusions yet (my Editor'd kill me) but suffice to say, I find the intersection of blogging, search, and the business information market to be pretty darn interesting.
I'm certainly not alone. Moreover has created "Enterprise-Grade Weblog Search" - essentially, a zietgiest mining tool for corporations. One can imagine similar products from any of the RSS search engines, or even from the major marketing agencies of the world."

JRDF

Well, it's not an impressive name but after talking to the Jena and Sesame people it seems important to have a consistent binding to RDF written in Java.

JRDF is going to be based on the best bits from Sesame, Jena, Kowari and RDF API Draft. Any other contributions would be good. Currently I've got the start of Blank Node, Graph, Literal, Node, NodeFactory, Statement and URIReference.

One of the annoying things is that the W3C specs for RDF don't talk about models anymore but graphs. It will be odd for a Model to implement Graph - maybe. Or there might be a lot of renaming to be done.

I've removed our implementation from Kowari and started using it instead. Still lots to do. Once I have Kowari done then it's Jena's turn.

Java Code

Abstract classes are not types "Have you ever seen code that declared a variable of type java.util.AbstractList? No? Why not? It's there, along with HashMap and TreeSet, etc. Because AbstractList is not a type. List is. AbstractList provides a convenient base from which to implement custom Lists. But I still have the option of implementing my own from scratch if I so choose. Maybe because I need some behaviour for my list such as lazy loading or dynamic expansion, etc. that wouldn't be satisfied by the default implementation."

"cglib is a powerful, high performance and quality Code Generation Library, It is used to extend JAVA classes and implements interfaces at runtime."

Tuesday, December 02, 2003

Blank Nodes

In RDF there are nodes. There are nodes that are resources (with or without URIs) and literals. The ones without URIs are blank nodes. These blank nodes can either be given a name (nodeID) or not. Statements are made of a subject (resource), predicate (resources with URIs) and object (anything). Simple enough.

Now I've been looking across three separate Java implementations. I found 8 implementations of classes designed to use blank nodes (resources without URIs):
* Two in Kowari,
* BNode, BNodeImpl and BNodeNode.
* AResource, Node_Blank and RDFNode.

What is maddening is that they each (well 7 of them) have a different way to get their name. What's even more maddening is this is right. Getting their name isn't a part of the RDF model - the most all of these different blank node implementations probably should have in common is equality and being the same type (just share a marker interface like Serializable).

Harpers.org brought to you by the Semantic Web

A New Website for Harper's Magazine "We cut up the Weekly Review into individual events (6000 of them, going back to the year 2000), and tagged them by date, using XML and a bit of programming. We did the same with the Harper's Index, except instead of events, we marked things up as “facts.”

Then we added links inside the events and facts to items in the taxonomy. Magic occured: on the Satan page, for instance, is a list of all the events and facts related to Satan, sorted by time. Where do these facts come from? From the Weekly Review and the Index. On the opposite side, as you read the Weekly Review in its narrative form, all of the links in the site's content take you to timelines. Take a look at a recent Harper's Index and click around a bit—you'll see what I mean.

The best way to think about this is as a remix: the taxonomy is an automated remix of the narrative content on the site, except instead of chopping up a ballad to turn it into house music, we're turning narrative content into an annotated timeline. The content doesn't change, just the way it's presented."

"A small team of Java coders and I are planning to take the work done on Harper's, and in other places like Rhetorical Device, and create an open-sourced content management system based on RDF storage. This will allow much larger content bases (the current system will start to get gimpy at around 30 megs of XML content—fine for Harper's, but not for larger sites), and for different kinds of content to be merged."

I'll have to look at Samizdat. Which "is a generic RDF-based engine for building collaboration and open publishing web sites." Seems to be the way things are heading.

Thursday, November 27, 2003

Simulated OS for teaching Assembly

Apoo is very similar to one of the uses of RCOSjava.

Jena2 Manager

"Briefly stated, I needed a means by which I could quickly hack models and ontologies to learn how to use Jena2. There remain many things to learn, and many things to finish coding in the program. I'm turning it loose so that others can contribute to its development. The JOSL license requires that those who fix things in the code or otherwise improve it return their code to the public. JOSL does not require that users use an open source license on new code that extends the licensed code."

Jena2 Manager

On Ontologies and Gnomes

The AI gnomes of Zurich "McDermott ends with a zinger:

It's annoying that Shirky indulges in the usual practice of blaming AI for every attempt by someone to tackle a very hard problem. The image, I suppose, is of AI gnomes huddled in Zurich plotting the next attempt to --- what? inflict hype on the world? AI tantalizes people all by itself; no gnomes are required. Researchers in the field try as hard as they can to work on narrow problems, with technical definitions. Reading papers by AI people can be a pretty boring experience. Nonetheless, journalists, military funding agencies, and recently the World-Wide Web Consortium, are routinely gripped by visions of what computers should be able to do with just a tiny advance beyond today's technology, and off we go again. Perhaps Mr. Shirky has a proposal for stopping such visions from sweeping through the population."

The entry links to a paper which lists the things that the Semantic Web "violates" wirth respect to traditional assumptions about AI. Including lack of referential integrity, variety in quality, diversity and no single authority. As noted, these are the same problems with human intelligence too.

Wednesday, November 26, 2003

webMethods going Semantic

Interview: webMethods CEO eyes Web services innovation "Secondly, there’s a whole other layer to deal with, what I call the semantic integration problem. Web services are great but they standardize pure connectivity between applications. The applications still have highly varied data models, extremely different ideas of what business processes should look like. Yet for most large organizations, a business process is going to span many applications. So you’re always going to need in the middleware stack something that can do wrapping, transformation, and, more than that, can actually keep the model of how the business processes are implemented across all of the infrastructure pieces.So [you need] something that’s technology-neutral underneath, like our Fabric product, and then on top have the ability to orchestrate business processes across all of these nodes in the fabric. Our customers now want to get real-time intelligence about what’s happening with the business and with the business processes, and they want to see it in dashboards, they want alerts. So we can put real-time monitoring around [IT infrastructure] at the business process level."

"We’re also able to offer enterprise event management, [injecting] business events into some kind of AI [artificial intelligence]-based rules engine."

Web Services in RDF

The question of how Web Services and the Semantic Web came up again recently. Here are a few links to current work in the area:
* Semantic Web enabled Web Services,
* Government Semantic XML Web Services Community of Practice,
* BPEL2DAML-S,
* Esperanto,
* Meteor-S,
* Supercharging WSDL with RDF, and
* SWAD-Europe Thesaurus Activity.

Tuesday, November 25, 2003

Don't Panic

Neil Gaiman hitchhikes through Douglas Adams' hilarious galaxy

Only 3

Three Uses for the Semantic Web They include:
* Sideline Semantics (or how to cut down on those darn post-planning columns),
* The Policy Ontology, and
* Cross Domain Searching for Calendar Concerns.

The Policy Ontology was perhaps the most interesting:
"I decided to tackle this with an interface to Jena for Apache Cocoon, or to use Cocoon parlance, a Jena-based transformer. I had no idea what kind of systems sat behind virtual reference applications, but I did know the protocol used underneath the queries was based on SOAP, and Cocoon excels at inserting itself in between any XML stream and adding value to the contents. So my approach was to use Jena's inference capabilities to map different classification schemes based on relationships defined in either RDF Schema or OWL. Yes, you could do the same thing with a table or two, and a thousand other ways, but the ontology approach provides a formal syntax for defining relationships."

Seems similar in idea to Sherpa Calendar.

The application is WIBS.

Fractally Yours

Openness & Interconnection "Big Fractal Tangle would be the name of a blog...A decade into the first Web, we've now got way too much information available, too much for any of us to sift through easily, which is why we need Round Two: the annotated, interconnected, Web. This new organic, evolving, maintainable, improvement will do more than simply increase the accuracy of our Google searches. It'll help real people understand and visualize interconnection, which in my opinion will alter our society profoundly for the better.

This was to be the point of my paper. Driving home that night, my brain frazzled and my voice hoarse from too much talk, I realized the topic was too big for a single paper. It'll have to be a blog."

See also: The Fractal nature of the Web.

Sunday, November 23, 2003

Random RDF Tools

MusicBrainz Java API, RDFical (English version), MnM and FOAF Explorer - some of these I've covered before.

One Stop Schema Shop

SchemaWeb " SchemaWeb is a repository for RDF schemas expressed in the RDFS, OWL and DAML+OIL schema languages.
SchemaWeb is a place for developers and designers working with RDF. It provides a comprehensive directory of RDF schemas to be browsed and searched by human agents and also an extensive set of web services to be used by RDF agents and reasoning software applications that wish to obtain real-time schema information whilst processing RDF data. "

Who are you?

Gillmor Takes On Dvorak's Anti-Blog Stance ""Perseus thinks that most blogs have an audience of about 12 readers," Dvorak argues. Yes, John, but who are those 12? If one of them is Bill Gates, and another is Tony Scott, CTO of General Motors, and another is John Cleese, well you get the idea. Sometimes it's who you know as much as what. RSS only amplifies this, allowing a Ray Ozzie to post only when it's valuable to him and his readers. It's "You've got blog.""

Bill, Tony, John, give me some feedback then.