Monday, November 29, 2004

IBM boosts Oncology Ontology

IBM and Massachusetts General Hospital Announce Effort to Improve Information Sharing Among Cancer Researchers ""Effective tools for information management, integrated tightly with underlying computing and data infrastructures, are key to life sciences researchers gaining new insights into complex problems," said David Grossman, Distinguished Engineer, IBM Internet Technology Group. "In addition, the use of semantic web technologies to integrate many sources and formats of data with advanced modeling algorithms is particularly helpful for this type of large-scale collaborative project.""

""There is an urgent need to develop a common, unifying infrastructure that enables the integration and sharing of knowledge about cancer -- both in terms of disparate data and distinct computational tools -- with the goal of modeling cancer as a complex dynamic system," said Dr. Deisboeck. "While advances in cancer research and new technologies have generated a wealth of new data and insight, all too often the lack of shared systems and standards makes integration of this crucial knowledge difficult or impossible.""

Python and PHP

The Next Language "The vast majority of J2EE deployments (over 80% according to Gartner) are simply Servlet/JSP to JDBC applications. Basically HTML front-ends to relational databases. It is ironic that much of what makes Java complicated today is all of its numerous band-aid extensions, such as generics and JSP templates, which were added to make these types of simple applications easier to develop."

"Apparently what is needed is a language/environment that is loosely typed in order to encapsulate XML well and that can efficiently process text. It should be very well suited for specifying control flow. And it should be a thin veneer over the operating system."

Just to make this clear, the idea that the future of application development is turning them into "a big text pump" seems rather foreign and completely opposite to where application development seems to be going.

Requirements and Architecture

Requirements guru shares 'cosmic truths' "Wiegers' list of cosmic truths also includes:

* "Customer involvement is the most critical factor in achieving software quality."
* "The customer is not always right, but the customer always has a point."
* "Change happens."
* "If it’s not in the requirements specifications, don’t expect to find it in the product."
* "Even the best requirements document cannot replace human dialog."
* "You are never going to have perfect requirements.""

Architected RAD gets an A in Gartner study "The Gartner survey of development teams, completed in the past month, found the ARAD approach reduces training time and increases productivity of coders regardless of the vendor tool used.

"We have gotten consistently positive feedback from users of Computer Associates' Advantage:Plex, Compuware's OptimalJ and IBM's Rational Rapid Developer offerings," writes Michael Blechar, one of the Gartner analysts who worked on the survey."

"Of the newest tool technology, the Gartner report says: ARAD methods and tools are just beginning to achieve recognition by mainstream Java 2 Platform, Enterprise Edition (J2EE) and .NET developers. The tools provide development teams with pre-built J2EE and .NET frameworks as well as pre-built technical components, which Gartner says can be customized by technical architects and used to generate 60 to 85 percent of the code. Then the programmers on the development team can add the business logic specific to the application."

Sunday, November 28, 2004

Learning with the Semantic Web

Reasoning and Ontologies for Personalized E-Learning in the Semantic Web "Adaptive educational hypermedia systems are able to adapt various visible aspects of the hypermedia systems to the individual requirements of the learners and are very promising tools in the area of e-Learning: Especially in the area of eLearning it is important to take the different needs of learners into account in order to propose learning goals, learning paths, help students in orienting in the e-Learning systems and support them during their learning progress...We propose a framework for such adaptive or personalized educational hypermedia systems for the semantic web. The aim of this approach is to facilitate the development of an adaptive web as envisioned e.g. in (Brusilovsky and Maybury, 2002). In particular, we show how rules can be enabled to reason over distributed information resources in order to dynamically derive hypertext relations. On the web, information can be found in various resources (e.g. documents), in annotation of these resources (like RDF-annotations on the documents themselves), in metadata files (like RDF descriptions), or in ontologies. Based on these sources of information we can think of functionality allowing us to derive new relations between information. "

Friday, November 26, 2004

Free trade that isn't free

Patently yours "We quite understand that (the title) How to Kill a Country may sound alarmist...We use the parallel experience of Canada to buttress some of these points. Canada is now being described by leading author, Mel Hurtig, as a "Vanishing country"...By the mid 1980s, about half of the major US corporations in Canada were 100-percent American-owned. Ten years later, some 85 per cent had no Canadian shareholders...As Canadian shareholders were eliminated, corporate boards were substantially reduced in size and more American directors were added, as were more U.S. CEOs and board chairmen. As external directors were eliminated, there was no longer a force to influence policy decisions which would be beneficial to Canada. Gone too was the ability to scrutinise the payment of dividends, management fees, and content costs paid to the parent company."

"But the Australian negotiators overlooked the point that Australia is a net importer of IPRs...As a whole, Australian industry has everything to gain by moving away from the Microsoft stranglehold and towards an Open Source mode - rather like governments in Germany and Taiwan are currently doing in earnest...local firms would do well to shift towards the Open Source model, and utilise open source programs such as Linux..."

"...frequently the actions are entirely justified, and entirely in the spirit of competition - as when an importer of copyright-protected CDs seeks them out in a third market and imports them, entirely legally, at a lower cost than is stipulated by the IPR-holder. The FTA makes this action much more difficult - in the name of placing severe restrictions on parallel imports. Another name for this is placing restrictions on free trade in IPR-protected goods - all within a "free trade" agreement!"

Dangling Databases

Why Relational Databases And Semantics Don't Mix "Jarg Corporation, which takes its name from "jargon", is about the next evolution of search. Actually, it is about more than that since semantics has wider applicability than search, but we will stick with search as an easily understood example of this technology."

"So, the question is: how does it do that?

The first answer is that it doesn't - at least in any general sense - only where it has already built an ontology which, in this case, is within healthcare. Indeed, its first customer is a hospital medical library. However, industry knowledge bases are becoming widely available and Jarg reckons that about 75% of the work involved in creating an ontology can be automated, so extending its product for new customers should not be a big issue."

"The second answer is that it achieves this sort of performance by refusing to use a relational database. Instead, it has patented its own approach, which involves storing semantic fragments. A segment fragment is either two elements and the relationship that joins them (for example, "Waterloo is a station") or it can store and element with a "dangling" relationship. This latter concept is especially important. The whole point about searching is that you want to be able to discover relationships that you didn't know existed."

MEST Architecture

The MEST architectural style "We have both agreed that we shouldn’t call our architectural style ProcessMessage after all. Instead, we decided to call it MEST (MESsage Transfer) so as to recognise the big influence REST had in this work and our thinking in general. So, after all the blog entries and the discussions with the community, we have finally arrived to the MEST architectural style of which ProcessMessage is part."

WWW2005 Tutorial: Architecting and Developing Message-Oriented Web Services "Savas and I have been accepted to present a tutorial at WWW2005 in Chiba next year. We're going to be talking about message-orientation and the MEST architectural style. Our approach is going to be very interactive: We'll be doing head-to-head live coding and will have the audience involved right the way through.

Broadly speaking, we're going to introduce a simple problem domain (probably a simple game), get the audience to work through the domain with us, identifying the services and message exchanges involved, then we'll code up a solution. Once we've got a solution in place we'll break it in various interesting ways and show how various WS-* protocols can help prevent such breakages from occuring."

RDF on Lambda

RDF and Databases "Some RDF research dropped me to a nice paper (PDF) from IBM discussing RDF with relational databases. This combination can replace half-baked application data mechanisms. These crop up regularly in my consulting work. Think nested directories of Windows INI files and brittle, binary files breaking on minor design iterations. The pain, the pain."

"There are several projects in this domain. My favorite so far is OpenRDF Sesame. It supports querying at the semantic level. It seems more mature than others, having derived from previous efforts, and works with both PostgreSQL and MySQL as well as Oracle. An abstraction layer called SAIL makes Sesame database-agnostic. Sesame even sports a stand-alone b-tree system, or in-memory operation, if you don't want an external database."

Thursday, November 25, 2004

One step closer to making a lighter Kowari

An initial port of Sesame 1.1's RIO RDF/XML parser to JRDF has been checked into JRDF's CVS repository. It's basically the same with a few modifications to the constructor, the SAXFactory explicitly asks for a reader that will use namespaces and a reduction on depending on other Sesame classes.

In Kowari it already uses JRDF to do N3 and RDF/XML exporting. A recent requirement that I was just asked about was providing RDF/XML from the result of an iTQL query. The client side JRDF API allows the creation of a JRDF graph using an iTQL answer so it might be possible to plug this into the exporter classes.

Instruction on getting it are here.

A simple example of using it:

public class RdfXmlParserExample {
public static void main(String[] args) throws Exception {
String baseURI = "http://slashdot.org/index.rss";
URL url = new URL(baseURI);
InputStream is = url.openStream();
final Graph jrdfMem = new GraphImpl();
RdfXmlParser parser = new RdfXmlParser(jrdfMem);
parser.setStatementHandler(new StatementHandler() {
public void handleStatement(SubjectNode subject,
PredicateNode predicate, ObjectNode object) {
try {
jrdfMem.add(subject, predicate, object);
}
catch (Exception e) {
e.printStackTrace();
}
}
}
);

parser.parse(is, baseURI);
Iterator iter = jrdfMem.find(null, null, null);
while (iter.hasNext()) {
System.err.println("Graph: " + iter.next());
}
is.close();
}
}


While mentioning Kowari, one of the included resolvers to be included in 1.1 will generate statements based on the latitude and longitude of two points. We used the DAML Geofile linked from Semantic Web Application Integration: Travel Tools.

Tuesday, November 23, 2004

15 parts of Classification Theory

Another JOT link, this time the 15 parts on "The Theory of Classification". Here they are: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, and 15.

Creating Requirements

Generating Complete, Unambiguous, and Verifiable Requirements from Stories, Scenarios, and Use Cases "Although very valuable as requirements elicitation, analysis, and initial validation tools, stories, scenarios, and use case path specifications are typically inadequate for specifying requirements because they are incomplete, ambiguous, and therefore unverifiable. For example, they usually do not address preconditions and postconditions, which have a huge influence on the meaning of the requirements. Similarly, they do not tend to state the triggering events that cause them to be true. They also do not typically clarify the distinctions between requirements (i.e., what the system must do and what postconditions it must ensure) and ancillary information (e.g., triggering events produced by actors and preconditions that may or may not be ensured). This column has provided examples and guidance on how to transform stories, scenarios, and use case path specifications into complete, unambiguous, and verifiable textual requirements."

Anti-metadata Google

Dear Mr. Bosworth "I can't believe that a smart guy like you could advocate Poscasting as a solution to the complexity of the semantic web and all that enterprise, corporate complexity - without a hint of meta-data - anywhere.

How's the work? Inference? Osmosis?

I know that Google is known as an anti-meta-data sort of place - but PLEASE oH LORd - get over that!

This is NOT about the religious wars of RSS 1.0 vs RSS 2.0. I couldn't give a dam about rdf, T B-L or any of that semantic web hooey.

I'd just like to see folks standardize on attributes, properties and meta-data around these new, burgeoning forms of micro-content."

Monday, November 22, 2004

Third SIGSEMIS

Volume 1, Issue 3 (PDF) "For the third time AIS SIGSEMIS bulletin is in your hands. Many interesting articles, a featured interview with Tom Gruber, our regular columns as well as several interesting announcements are waiting for your attention."

Interview with Tom Gruber: "In fact, the World Wide Web is based on a semiformal ontology, and it shows how ontological commitment works in software interoperability. At its core, the concept of the hyperlink is based on an ontological commitment to object identity. In order to hyperlink to an object requires that there be a stable notion of object and that its identity doesn’t depend on context (which page I am on now, or time, or who I am). Most of the machinery of the early Web standards are specifications of what can be an object with identity, and how to identify it independently of context. These standards documents serve as ontologies – specifications of the concepts that you need to commit to if you want to play fairly on the Web. If one built a system with these commitments, all of the web infrastructure works well."

"Intraspect was designed on the assumption that it is more valuable to get evidence of human knowledge into a collective memory than to add structure to existing online material. So we created technology that helped people work together on-line, and as a byproduct their work became available for discovery using information retrieval technology."

Other interesting articles: "Component Requirements for a Universal Semantic Web Framework", "Elements of a First Visual Rule Language for the Semantic Web" (about REWERSE), "Response Management in Multidimensional Web Information Systems", "The Semantic Web Trends in Brief", "Using e-business Registry / Repository for E-Health Semantics" and many others.

Your blog is boring and other links

* RSS, Blogs on a Roll, But How Extensive is Their Use? and a response.
* Swebok All you need to know about Software Engineering.
* Best Software Essays of 2004 via Danny.
* ISCOC04 Talk on how RSS 1.0 failed, with lots of other comments that seem to support a RESTful like system like RDF. With followups Fielding Bosworth and Quick Reactions.
* The Many Faces of J2EE, v5.0 Another annotations for J2EE 1.5 piece as well as Google and JBoss join SE/EE Java Council.
* Domain Speific Language and Domain Specific Modelling. Is UML really the best tool for the job?. The end of UML?

Friday, November 19, 2004

Metadata does Matter

In a similar vein to, I.T. does Matter a recent column entitled Does Meta Data Matter? highlights the competitive advantage that metadata plays in the enterprise.

"Is meta data strategic in nature? Yes. The reality is that information technology continues to get more complex. Our ability to manage these technologies and solutions requires a higher degree of knowledge and management skills. In the dynamic environment we see emerging, command and control style of management fails to deliver a competitive advantage. As a wise man once said, "All great things have been in done in spite of management." Our ability to adapt within the technology community may be dependent on our ability to handle multiple tasks, objectives and strategies which can then change on a dime. Meta data plays a central roll in your organization's ability to become an agile organization. Moving to common infrastructures, software platforms and even systems does not negate the competitive advantages that technology and meta data can bring."

Keeping abreast of the Semantic Web

Revolutionising breast cancer treatment through knowledge management "The system uses Semantic web technologies, enabling information from X-ray mammograms, MRI images, biopsy results and data from the clinician to be made available when the practitioners meet for their weekly Triple Assessment Procedure. Semantic web technologies allow information to be linked in such a way that it can be easily processed by machines. Practitioners can then view different types of images and scans, call up patient information, and automatically generate reports. It is also possible to investigate, annotate and analyse the data using web and Grid services.

‘This research draws on technologies in which the UK is a world leader,’ says Professor Nigel Shadbolt of the School of Electronics and Computer Science at the University of Southampton. ‘Eventually, e-health will be delivered using the web and incredibly powerful networks of computers. Medical practitioners will have the information and evidence at their fingertips to support decision-making that has a direct impact on us all.’"

More Working Notes

Andrae is out doing both Paul and I with two blogs: Etymon which has links to papers, interesting articles and the like and Circumlocute which is similar to Working notes.

When two worlds combine

"Gnowsis and Fenfire have met and another time, great Semantic Web developers (aka us) have proven that using ontologies, RDF and web protocols RULEZ. In just a day, two open source projects made substantial integration work."

An early screenshot of Fenfire and Gnowis combining their efforts.

Google Scholar

Hits for the Semantic Web gives you that piece in Scientific American. Another interesting one is citations to "The Description Logic Handbook: Theory, Implementation and Applications".

Sesame 1.1

Sesame 1.1 released "Highlights of this release are:

* The Graph API, an extension of Sesame's access APIs, allows fine-grained manipulation of RDF models directly from Java.
* The Native Disk Store is a new storage backend that works directly on the file system, without need for a DBMS. It uses B-Tree indexing on binary files for fast, efficient and scalable storage.
* SeRQL revision 1.1 is a syntax revision that makes SeRQL queries even easier to read and write, and makes embedding in XML easier.
* Blank node handling has dramatically improved compared to 1.0.x.
* Lots of issues related to full Unicode support have been fixed.
* RDF Schema inferencing has been updated to be fully compliant with the W3C RDF Semantics Recommendation.
* Support for MS SQL Server as storage backend RDBMS. Thanks to Adam Skutt for providing fixes and suggestions for this.
* The Rio parser now supports the Turtle serialization format.
* Partial OWL reasoning support through Sesame's custom inferencer.
* Fully updated and extended User Documentation, including code examples for use of the Sesame APIs and a new Troubleshooting and FAQ chapter."