- Selenium IDE 0.7 Released "It is 100% functional by itself, or it can be used in combination with Selenium and any of the language (Java, Ruby, etc) drivers, provided you are willing to write the output template."
- JSONC, JSONI, JSONP: Creating tailored JSON from SPARQL query results "What's missing (if we want to avoid custom, server-side code or heavy pre-processing on the client) is a way to tell the SPARQL endpoint to arrange and index the tabular results before they are serialised as JSON: JSONI. The jsoni parameter works similar to the jsonc one, but it allows nesting of result vars to specify index structures" - JSON (JavaScript Object Notation)
- Grady Booch links to "Why I Hate Frameworks": "According to our research, what people really needed wasn't a Universal Hammer after all. It's always better to have the right kind of hammer for the job. So, we started selling hammer factories, capable of producing whatever kind of hammers you might be interested in using."
- On Rigid Rules "The getters and setters on their own aren't a problem. It's how and why they are called that can cause problems. If other objects are calling getters and setters then doing work on the result, it's likely that many objects need to do the same work. This results in duplicated code. Duplicated code causes rigidity. Any design change that affects this code will affect all
copies requiring changes to be made across the system. It's better to have one copy of the code performing this operation and the natural place to put that code is in the original object." - USB could sap Core Duo's power "Tom’s Hardware Guide, has discovered a glitch in one Core Duo/Napa system that suggests machines based on the platform might see unexpectedly short battery life, when operated on battery power with USB 2.0 devices are attached to them."
- Special Topics B Proposal "...SET is designed specifically for commerce [allowing it to] have a 1024-bit cipher block [Vis96b]. Combining this with the backing of major Internet and credit card companies, SET is destined to become the standard means of transmitting electronic, commercial data in the next year or so." And for what really happened.
- SWEO Suggestion #2 : Semantic Web in a Box Lists the key properties of a turn-key Semantic Web application. I would add some maybe metadata extraction and P2P ability but that's probably too much bloat.
Monday, February 06, 2006
New URLs
Saturday, February 04, 2006
Flash Google
Two Flash version of Google's fine work:
All part of AFLAX.
- Google Earth - FlashEarth.
- Google Talk, GMail and Flickr - Gtalkr.
- Google News - Newsmap
All part of AFLAX.
Thursday, February 02, 2006
Dueling Seminal Papers
So, *ahem* in "No Silver Bullet: Essence and Accidents of Software Engineering" Fred Brooks writes: "I think the most important single effort we can mount is to develop ways to grow great designers...Great designers and great managers are both very rare...each software organization must determine and proclaim that great designers are as important to its success as great managers are, and that they can be expected to be similarly nurtured and rewarded."
This basically looks like creating a corporate hierarchy of designers similar to corporate executives.
Contrast this with, "Necessary Preconditions for the Bazaar Style" part of "The Cathedral and the Bazaar" where Eric Raymond writes, "I think it is not critical that the coordinator be able to originate designs of exceptional brilliance, but it is absolutely critical that the coordinator be able to recognize good design ideas from others...Linus, while not (as previously discussed) a spectacularly original designer, has displayed a powerful knack for recognizing good design and integrating it into the Linux kernel."
From my experience new ideas aren't successfully applied from the top down but rather from the bottom up. By relying on a few to understand new ideas or to implement them in fully formed designs actually encourages stagnation. It is frequently too risky and expensive to try a new design, applied in a top down manner across an entire project or framework. The better solution seems to be to grow design and better design processes, not designers.
This basically looks like creating a corporate hierarchy of designers similar to corporate executives.
Contrast this with, "Necessary Preconditions for the Bazaar Style" part of "The Cathedral and the Bazaar" where Eric Raymond writes, "I think it is not critical that the coordinator be able to originate designs of exceptional brilliance, but it is absolutely critical that the coordinator be able to recognize good design ideas from others...Linus, while not (as previously discussed) a spectacularly original designer, has displayed a powerful knack for recognizing good design and integrating it into the Linux kernel."
From my experience new ideas aren't successfully applied from the top down but rather from the bottom up. By relying on a few to understand new ideas or to implement them in fully formed designs actually encourages stagnation. It is frequently too risky and expensive to try a new design, applied in a top down manner across an entire project or framework. The better solution seems to be to grow design and better design processes, not designers.
Hypercode - Beyond AJAX
There's nothing like combining two technologies together to produce something that sounds cool, Redfoot, "...can be best described as a webized Operating System. An operating system built around the notion of hypercode and where the persistence is built around the notion of an RDF Graph rather than a File Tree (see the namesys Graph vs. Tree). It provides standard web mechanisms to transport information across different machines and programmatically specifies the tasks for bundling and installing new software features across networked machines. The functions of Redfoot are analogous to that of an operating system kernel that manages resources across the web. To allow humans and other HTTP service agents to utilize Redfoot, it uses RDFLib and third party libraries such as Twisted and Kid to convert between RDF-based information and HTML or other popular formats. Therefore, each Redfoot server, is designed to interacting with a human, other redfeet, or any other HTTP-aware computing service. The design and implementation principles of Redfoot is to recursively utilize existing standards and libraries whenever and whereever it can, resulting in a highly compact, yet functionally rich execution kernel."
Hypercode "...is mainly designed as a portable database that helps both programmers and machines to exchange either executable or structured data...specifically utilizes RDF to capture both the structure and meaning of raw data in the XML-based global naming context, therefore, it has the potential to better utilize information in the code's execution context when necessary. In other words, Hypercode dosen't only treat Python or other scripting language as textual data, the defacto interpreter, Redfoot server, could adaptively leverage the supporting resources and execution policies defined in the associated RDF content to control the rights and behavior of script execution."
I was looking at O'Reilly's Twisted book this morning.
I don't know where I read this but it seems cool enough to download and try - like the Bookmark aspect.
Hypercode "...is mainly designed as a portable database that helps both programmers and machines to exchange either executable or structured data...specifically utilizes RDF to capture both the structure and meaning of raw data in the XML-based global naming context, therefore, it has the potential to better utilize information in the code's execution context when necessary. In other words, Hypercode dosen't only treat Python or other scripting language as textual data, the defacto interpreter, Redfoot server, could adaptively leverage the supporting resources and execution policies defined in the associated RDF content to control the rights and behavior of script execution."
I was looking at O'Reilly's Twisted book this morning.
I don't know where I read this but it seems cool enough to download and try - like the Bookmark aspect.
Wednesday, February 01, 2006
Get Functional
Hadn't used Commons Collections until recently, as Tom recently blogged about Closure Time for Java.
It includes Predicates for equals, and and others.
Related: Functional programming in the Java language and Commons Primitives package (another implementation of primitive collections).
It includes Predicates for equals, and and others.
Related: Functional programming in the Java language and Commons Primitives package (another implementation of primitive collections).
Tuesday, January 31, 2006
Visiting Python and Smalltalk
Something that I continually pester people about, after working on a message passing system in my very first big Java project, the switch statement is not OO, "The switch statement has all sorts of problems with cohesion and coupling. The irony today is that almost all OO languages today, other than Python and Smalltalk (in which it can easily be implemented), still have switch statements."
Also, OO - the Solution to the Problem of Switch Statements which links to Replacing Conditional with Polymorphism.
Also, OO - the Solution to the Problem of Switch Statements which links to Replacing Conditional with Polymorphism.
Sunday, January 29, 2006
Beware the Chrysanthemum Men
Can Corporations Stop Doing Evil? "Aggressive and competitive tendencies assisted our ancestors to survive. That's why they are there. Effective channeling, not avoidance, seems necessary. We can do that with diversion ("Yay! The Steelers won! Woo hoo!") or direction (such as the very existence of the Special Forces and the CIA), but we had better not ignore it."
This posting reminded me of a few things that I had intended to blog but hadn't.
Are we designed for violence? "Anyone doubting that treating other people as more than instruments is founded in the brain would do well to look into developments in the study of self–other mapping. This has provided stronger and stronger evidence that these relationships are hardwired into us, strikingly with the discovery of mirror neurons that fire in the same way for events that occur to you or to those you observe (Gallese and Goldman 1998). Many argue that empathy is an outcome of these representations (see e.g. Frith and Frith 1999). And recent research demonstrates appreciating someone else's pain activates many of the same areas as experiencing it (Jackson, Meltzoff, & Decety 2004): good evidence for a VIM [Violence Inhibition Mechanism] -like mechanism, and certainly a rebuttal to those who think our withdrawal from violence is unnatural...We are not innately disposed to violence, or even indifferent to violence, we are neurologically bound away from violence. This understanding gives us a solid basis for treatment, and an honest beginning from which to address the continuing problem of violence in society."
Mirror neurons seems to be a theory that is reinforcing the idea of memes but also changing it slightly.
A NY Times article has more details (PDF) including the effect this might have on children: "Mirror neurons work best in real life, when people are face to face. Virtual reality and videos are shadowy substitutes...a study in the January 2006 issue of Media Psychology found that when children watched violent television programs, mirror neurons, as well as several brain regions involved in aggression were activated, increasing the probability that the children would behave violently."
This idea also has influenced other areas of research including: language, freedom of speech and autism.
Also, related is The Milgram experiment (as mentioned in "Enron: The Smartest Guys in the Room") and an interview last year with Peter Cundall showing that the capacity to do evil is everywhere:
"ANDREW DENTON: How do you break a human being down to a point of amorality?
PETER CUNDALL: By giving them power. By giving them absolute power over people. I suppose I see, today, the biggest problem in society is what we call the "control freaks". These are the malignant narcissists, right. The people who are - and we've met them. Look, you can go to a tiny organisation, a school parents and friends, a trades union, ABC, we know who the control freaks - it's true, we know them, we know who they are and they want - that's right.
ANDREW DENTON: We've just been cut off.
PETER CUNDALL: But they're people without a gone conscience and without any pity, and there is not many of them. But when you and I are fast asleep at night, they're lying awake, scheming, and they're hard to compete with, and I'm not joking. You know - anybody that's worked in an office or anything, there is always someone that wants to take control, right. The supreme example, of course, is people like Stalin and Hitler. Hitler, don't forget, almost his last days, when he was in the bunker, he was saying at one time that, "The SS have betrayed me," because they retreated, and he had all his commanders executed. And almost his last words was, "The German people have betrayed me," you know. I mean - and this is an example of a supreme form of pathological narcissism and you do get it. You get it in politics. They're the people who can't admit that they have made a mistake.
ANDREW DENTON: Do you get it in gardening?
PETER CUNDALL: Absolutely. You get it in gardens and in gardening, right. You get somebody - mainly with blokes. Blokes and women are totally different gardeners. Blokes are single-minded, right. So you get these dahlia men, right - really, it's true - and they compete with each other, or the cactus men and they compete, right, and the chrysanthemum men."
This posting reminded me of a few things that I had intended to blog but hadn't.
Are we designed for violence? "Anyone doubting that treating other people as more than instruments is founded in the brain would do well to look into developments in the study of self–other mapping. This has provided stronger and stronger evidence that these relationships are hardwired into us, strikingly with the discovery of mirror neurons that fire in the same way for events that occur to you or to those you observe (Gallese and Goldman 1998). Many argue that empathy is an outcome of these representations (see e.g. Frith and Frith 1999). And recent research demonstrates appreciating someone else's pain activates many of the same areas as experiencing it (Jackson, Meltzoff, & Decety 2004): good evidence for a VIM [Violence Inhibition Mechanism] -like mechanism, and certainly a rebuttal to those who think our withdrawal from violence is unnatural...We are not innately disposed to violence, or even indifferent to violence, we are neurologically bound away from violence. This understanding gives us a solid basis for treatment, and an honest beginning from which to address the continuing problem of violence in society."
Mirror neurons seems to be a theory that is reinforcing the idea of memes but also changing it slightly.
A NY Times article has more details (PDF) including the effect this might have on children: "Mirror neurons work best in real life, when people are face to face. Virtual reality and videos are shadowy substitutes...a study in the January 2006 issue of Media Psychology found that when children watched violent television programs, mirror neurons, as well as several brain regions involved in aggression were activated, increasing the probability that the children would behave violently."
This idea also has influenced other areas of research including: language, freedom of speech and autism.
Also, related is The Milgram experiment (as mentioned in "Enron: The Smartest Guys in the Room") and an interview last year with Peter Cundall showing that the capacity to do evil is everywhere:
"ANDREW DENTON: How do you break a human being down to a point of amorality?
PETER CUNDALL: By giving them power. By giving them absolute power over people. I suppose I see, today, the biggest problem in society is what we call the "control freaks". These are the malignant narcissists, right. The people who are - and we've met them. Look, you can go to a tiny organisation, a school parents and friends, a trades union, ABC, we know who the control freaks - it's true, we know them, we know who they are and they want - that's right.
ANDREW DENTON: We've just been cut off.
PETER CUNDALL: But they're people without a gone conscience and without any pity, and there is not many of them. But when you and I are fast asleep at night, they're lying awake, scheming, and they're hard to compete with, and I'm not joking. You know - anybody that's worked in an office or anything, there is always someone that wants to take control, right. The supreme example, of course, is people like Stalin and Hitler. Hitler, don't forget, almost his last days, when he was in the bunker, he was saying at one time that, "The SS have betrayed me," because they retreated, and he had all his commanders executed. And almost his last words was, "The German people have betrayed me," you know. I mean - and this is an example of a supreme form of pathological narcissism and you do get it. You get it in politics. They're the people who can't admit that they have made a mistake.
ANDREW DENTON: Do you get it in gardening?
PETER CUNDALL: Absolutely. You get it in gardens and in gardening, right. You get somebody - mainly with blokes. Blokes and women are totally different gardeners. Blokes are single-minded, right. So you get these dahlia men, right - really, it's true - and they compete with each other, or the cactus men and they compete, right, and the chrysanthemum men."
Wednesday, January 25, 2006
Build One to Throw Away
BenchMark-0.9.5 "Here are some Benchmark results between Wine and Windows XP, Your mileage will vary depending on your Linux config, Wine version and Hardware...Wine has the current lead on 67 tests"
Via, "Is Wine really faster than Windows?".
Via, "Is Wine really faster than Windows?".
Tuesday, January 24, 2006
The Best Code is Shared Code
Code Reviews: Just Do It (but really just a quote from Steve McConnell) "...software testing alone has limited effectiveness -- the average defect detection rate is only 25 percent for unit testing, 35 percent for function testing, and 45 percent for integration testing. In contrast, the average effectiveness of design and code inspections are 55 and 60 percent."
This is found in "Chapter 21: Collaborative Construction" of the 2nd edition of "Code Complete". Furthermore, he goes on to say that "The cost of full-up pair programming is probably higher than the cost of solo development - on the order of 10-25 percent higher - but the reduction in development time appears to be on the order of 45 percent..." The defect detection rate, stated in the book, is listed as 40-60%.
It also includes some tips for successful pair programming: coding standards, active, frequently rotated pairs, a good pairing environment, etc. Of course, he seems to think that pair programming involves one keyboard, one mouse and one monitor - which doesn't seem like a very good environment at all.
The hard data section contains lots of references to the benefits of early detection of defects, which are also mentioned in a paper, "Planning to Get the Most Out of Inspection". Similarly, "Testing is one of the Least Effective Defect Removal Techniques".
From a recent XP meeting.
This is found in "Chapter 21: Collaborative Construction" of the 2nd edition of "Code Complete". Furthermore, he goes on to say that "The cost of full-up pair programming is probably higher than the cost of solo development - on the order of 10-25 percent higher - but the reduction in development time appears to be on the order of 45 percent..." The defect detection rate, stated in the book, is listed as 40-60%.
It also includes some tips for successful pair programming: coding standards, active, frequently rotated pairs, a good pairing environment, etc. Of course, he seems to think that pair programming involves one keyboard, one mouse and one monitor - which doesn't seem like a very good environment at all.
The hard data section contains lots of references to the benefits of early detection of defects, which are also mentioned in a paper, "Planning to Get the Most Out of Inspection". Similarly, "Testing is one of the Least Effective Defect Removal Techniques".
From a recent XP meeting.
Monday, January 23, 2006
On the Other Hand
A collection of posts that seem contrary to my current point of view but which I find myself agreeing with none the less:
- Instead of autoboxing in Java being confusing, inconsistent and clunky it's useful and concise.
- Instead of Java being dead because of its strong typing people are increasingly picking more strongly typed languages and Perl is being replaced by Ruby.
- Instead of test driving making creating procedural/data driven code and to avoid any test artifacts, test driving leads to more OO code and methods for testing should be left in (although I would've thrown an IllegalArgumentException not a NPE).
- Instead of everyone becoming a librarian, librarians and semantic web developers are excluding many means of shared understanding.
Gnarly Client Code
Firebug "FireBug is a new tool for Firefox that aids with debugging Javascript, DHTML, and Ajax. It is like a combination of the Javascript Console, DOM Inspector, and a command line Javascript interpreter."
Includes the ever important "printfire" to send debug messages to the console.
Related to, December Java Users Group talk on AJAX.
Includes the ever important "printfire" to send debug messages to the console.
Related to, December Java Users Group talk on AJAX.
Saturday, January 21, 2006
A Post from the Future
Folktologies -- Beyond the Folksonomy vs. Ontology Distinction "One point that Clay makes, which I think is very interesting, is his view that perhaps the world is moving from a graph-theory information model (ontologies) to a set-theory model (folksonomies) -- but in fact, under the surface this argument falls apart. OWL is nothing other than a language for enabling extremely sophisticated set-theoretic operations on information. In fact, if you actually look at the OWL language itself, it is primarily comprised of set-theoretic statements."
"Ontologies are, in my opinion, simply the next evolution of database schemas. Surely, Clay would not argue that database schemas have no place in the world!"
"Imagine a folksonomy combined with an ontology -- a "folktology." In a folktology, users could instantly propose or modify ontological classes and properties in the same manner that they do with tags in tagging systems. The most popular ontological constructs (the most-instantiated classes, or slots on classes, for example) would "rise to the top" and self-amplify, while the less-instantiated ones would "fall to the bottom" over time. In this way an emergent, self-organizing, and self-pruning ontology could emerge within a community."
Via, Folktologies -- Beyond the Folksonomy vs. Ontology Distinction.
"Ontologies are, in my opinion, simply the next evolution of database schemas. Surely, Clay would not argue that database schemas have no place in the world!"
"Imagine a folksonomy combined with an ontology -- a "folktology." In a folktology, users could instantly propose or modify ontological classes and properties in the same manner that they do with tags in tagging systems. The most popular ontological constructs (the most-instantiated classes, or slots on classes, for example) would "rise to the top" and self-amplify, while the less-instantiated ones would "fall to the bottom" over time. In this way an emergent, self-organizing, and self-pruning ontology could emerge within a community."
Via, Folktologies -- Beyond the Folksonomy vs. Ontology Distinction.
Mocking Correctly
A really great posting, "Best and Worst Practices for Mock Objects" lists a number of best practices for mocking.
A few more occurred to me:
The previous post in the series, "Why and When to Use Mock Objects" lists what makes a good unit test including: being atomic, order independent and isolated, intention revealing, easy to setup (a unit test code smell is a lot of setup code) and runs fast.
There also looks like a really interesting site, Test Automation Patterns. It lists test code smells including: test code duplication is bad, obscure tests, and data sensitivity. Also, a list of fixture strategies: standard, fresh and shared.
- "Be careful about mocking or stubbing any interface that is outside your own codebase. If you don’t truly understand the semantics and proper usage of an interface you don’t have any business mocking that interface in a unit test. My best advice is to create your own interface to wrap the interaction with the external API classes."
- "My advice is to just test the data access code against the actual database with integration tests. For the most part I think testing persistence code without the database is a complete waste of time."
- "I’ve worked with some people before that felt that there should never be more than 1-2 mock objects involved in any single unit test. I wouldn’t make a hard and fast rule on the limit, but anything more than 2 or 3 should probably make you question the design...Minimizing the number of mock objects necessary for any given unit test is a very effective way of keeping the Cyclomatic Complexity of your classes below an acceptable level."
- "Ideally you only want to mock the dependencies of the class being tested, not the dependencies of the dependencies. From hurtful experience, deviating from this practice will create unit test code that is very tightly coupled to the internal implementation of a class’s dependencies."
A few more occurred to me:
- Use as many patterns and idioms in your unit tests that you do in regular code, like testing exceptions and fixtures. This can be a great way to reduce the drudgery of testing bean objects for instance (not that I in anyway condone the use of bean objects).
- Allow reuse of commonly used objects between tests.
- The length and complexity of a unit tests is an indication of the complexity of the underlying object - once a test reaches a certain length/complexity it's time to refactor the object being tested.
- Duplication in tests is just as bad as duplication in the code being tested.
- No code goes into production just for the sake of testing.
- Use reflection to gain access to internal object state. This is related to the previous point as a way to avoid artifacts of testing in production code. Other ideas include a good equals method, package level access and getter methods - but these leave code that may only exist for testing. And the last two are also a good way to break encapsulation.
- Something that is easy to test is not always good design. For instance, having a fixture or utility code that makes bean objects easy to test drive shouldn't dictate using bean objects for everything.
The previous post in the series, "Why and When to Use Mock Objects" lists what makes a good unit test including: being atomic, order independent and isolated, intention revealing, easy to setup (a unit test code smell is a lot of setup code) and runs fast.
There also looks like a really interesting site, Test Automation Patterns. It lists test code smells including: test code duplication is bad, obscure tests, and data sensitivity. Also, a list of fixture strategies: standard, fresh and shared.
Thursday, January 19, 2006
IntelliJ vs Eclipse Deathmatch
For my own reference, I'll put comparisons of Eclipse vs IntelliJ. I didn't like Eclipse because I found setting up projects a pain so it was a bit of a non-starter.
Tom's guide: "Eclipse requires you to save all the time...Eclipse has serious usability issues in graphical diff...Like Word, Eclipse wants to own your files and wants you to make all your changes through it".
It's subject of a recent Slashdot article.
It'd be good to find a comparison of the refactorings that are missing/different in Eclipse. I'm also trying out NetBeans 5 to see what it's like.
Update: An old comparison of Eclipse and IntelliJ and a PDF comparing RefactorIt, Eclipse 3.0 and IntelliJ 4.5. A list of new refactoring in NetBeans. Eclipse 3.2 M4 reproduces many features of Swing and IntelliJ IDEA, one of the threads Is $500 really alot of money?.
Tom's guide: "Eclipse requires you to save all the time...Eclipse has serious usability issues in graphical diff...Like Word, Eclipse wants to own your files and wants you to make all your changes through it".
It's subject of a recent Slashdot article.
It'd be good to find a comparison of the refactorings that are missing/different in Eclipse. I'm also trying out NetBeans 5 to see what it's like.
Update: An old comparison of Eclipse and IntelliJ and a PDF comparing RefactorIt, Eclipse 3.0 and IntelliJ 4.5. A list of new refactoring in NetBeans. Eclipse 3.2 M4 reproduces many features of Swing and IntelliJ IDEA, one of the threads Is $500 really alot of money?.
Tuesday, January 17, 2006
Definition of Open Source
"Software that takes longer to install and configure than it would to write yourself."
This isn't what I truely think, but after several frustrating experience with LAMP software (especially PHP) it feels like it.
This isn't what I truely think, but after several frustrating experience with LAMP software (especially PHP) it feels like it.
Monday, January 16, 2006
Query Anything
Access In-memory Objects Easily with JoSQL "Once in a while something comes along that is so simple, so straightforward, and so obvious that it's amazing that nobody did it long ago."
An example query,
I've blogged previously, Using RDF to improve Object-First Development. This is an example that is the same principal as Kowari's resolvers (or any idea of smushing metadata like Gnowsis) - allowing anything to be treated as RDF (or in JoSQL's case SQL).
An example query,
SELECT *
FROM java.io.File
WHERE name $LIKE "%.html"
AND lastModified BETWEEN toDate('01-12-2004')
AND toDate('31-12-2004')
I've blogged previously, Using RDF to improve Object-First Development. This is an example that is the same principal as Kowari's resolvers (or any idea of smushing metadata like Gnowsis) - allowing anything to be treated as RDF (or in JoSQL's case SQL).
DSLs, LOP and Ruby
Ruby on Rails vs. Java: An Expert Roundtable "...you can start doing that now in Java with language-oriented programming because there are some new tools coming out. Martin Fowler coined this term "language-oriented programming (LOP)." In fact, he's written a paper about it. The paper will be influential because new tools are coming out—tools that let you build a domain-specific language on top of the Java Machine. There's a product coming out (based on JetBrains's Meta Programming System) that I am showing here at No Fluff, Just Stuff.
Intentional Software (Charles Simonyi's company) is coming out with a tool soon. Microsoft is doing it with something called "software factories". People have written domain-specific languages in dynamic languages like Ruby, because a dynamic language is more suited for creating a DSL. But now we are starting to see tools that allow you to use a DSL in a strongly type language like Java."
"One of the things that delights me is seeing that testing is in Ruby on Rails from the first minute. Of course you want to build your application and test it, but with Ruby on Rails we all build our applications and test them the same way. So I know when I sit down on somebody else's Ruby on Rails project, it's not just that the language, the framework, the templates, and all those things are going to be the same. The discipline for writing the automated tests at different levels is going to look the same across projects, so I can jump in and immediately be productive. I feel like when I sit down to do a Java application today, or join some team that is working on one, I'll immediately have to figure out how their automotive testing is done—or worse yet I'll have to design the automated testing because they haven't done it yet."
Intentional Software (Charles Simonyi's company) is coming out with a tool soon. Microsoft is doing it with something called "software factories". People have written domain-specific languages in dynamic languages like Ruby, because a dynamic language is more suited for creating a DSL. But now we are starting to see tools that allow you to use a DSL in a strongly type language like Java."
"One of the things that delights me is seeing that testing is in Ruby on Rails from the first minute. Of course you want to build your application and test it, but with Ruby on Rails we all build our applications and test them the same way. So I know when I sit down on somebody else's Ruby on Rails project, it's not just that the language, the framework, the templates, and all those things are going to be the same. The discipline for writing the automated tests at different levels is going to look the same across projects, so I can jump in and immediately be productive. I feel like when I sit down to do a Java application today, or join some team that is working on one, I'll immediately have to figure out how their automotive testing is done—or worse yet I'll have to design the automated testing because they haven't done it yet."
Friday, January 13, 2006
Don't put all your Pasta in the One Basket
What Sort of Pasta Do You Want? "After reflecting on a system that was developed using techniques found more so in agilest development teams such as Test Driven Development (TDD) and dependency injection, they observed that the parts of the system were more loosely coupled and more easily interchangeable, good indicators that it would be a better system to maintain. The code is better described as ravioli code instead of that of its more common pasta brethren.
This analogy has really stuck with me since because of the number of parallels it draws. Take one such example – the reason that ravioli is typically more expensive than spaghetti, even though they are both made from the same fundamental ingredients, is that making good ravioli takes a lot more skill than it does spaghetti. This idea is,of course, not new, and can be taken to extremes (see Wikipedia’s entry) but I know which one is is my favourite."
So ravioli code is part of the pasta theory of programming which includes lasagne code.
A little bit of Spring, makes your code Sing! "We make extensive use of Spring at the day job, and feel it has been a key contributor to our extremely high code coverage, and the promotion of the ravioli code pattern; although some are of the opinion that ravioli code is an anti-pattern, I prefer smaller pieces loosely joined over the alternative."
It is an anti-pattern in the way that too many, loosely joined, tiny pieces of code hampers the understanding of the code. Too many classes with 10 lines in them calling another class, just catching an exception or looping around a method call, etc. In this way, it can reduce the ease at which code can be modified.
In "Where Smalltalk Went Wrong 2": "Again, ravioli code doesn't offer a path from here to there. "Here" in this case is the ignorant programmer first introduced to the code, "there" is the same programmer successfully making whatever change they need to make. In my experience with ravioli code there is a gestalt which is arrived at in a sudden realization, when the picture comes together. But before you reach that realization you find yourself in a murky set of code where it is difficult to predict effects or determine causality, and you seldom know how far away you are from the gestalt realization.".
To push the analogy, the ravoli becomes too small so that there is nothing left but empty shells of pasta. It's much more palatable, code-wise, than spaghetti (where the goodness is all over the plate and can't be stuffed in) or mud.
In the same way that there is duplication detection there should also be a metric available to judge whether a piece of code is not providing enough functionality or value. However, the way most code is written this isn't frequently a problem. A good example given of "about the right size" is where the algorithm is on paper and the code is in our head.
This analogy has really stuck with me since because of the number of parallels it draws. Take one such example – the reason that ravioli is typically more expensive than spaghetti, even though they are both made from the same fundamental ingredients, is that making good ravioli takes a lot more skill than it does spaghetti. This idea is,of course, not new, and can be taken to extremes (see Wikipedia’s entry) but I know which one is is my favourite."
So ravioli code is part of the pasta theory of programming which includes lasagne code.
A little bit of Spring, makes your code Sing! "We make extensive use of Spring at the day job, and feel it has been a key contributor to our extremely high code coverage, and the promotion of the ravioli code pattern; although some are of the opinion that ravioli code is an anti-pattern, I prefer smaller pieces loosely joined over the alternative."
It is an anti-pattern in the way that too many, loosely joined, tiny pieces of code hampers the understanding of the code. Too many classes with 10 lines in them calling another class, just catching an exception or looping around a method call, etc. In this way, it can reduce the ease at which code can be modified.
In "Where Smalltalk Went Wrong 2": "Again, ravioli code doesn't offer a path from here to there. "Here" in this case is the ignorant programmer first introduced to the code, "there" is the same programmer successfully making whatever change they need to make. In my experience with ravioli code there is a gestalt which is arrived at in a sudden realization, when the picture comes together. But before you reach that realization you find yourself in a murky set of code where it is difficult to predict effects or determine causality, and you seldom know how far away you are from the gestalt realization.".
To push the analogy, the ravoli becomes too small so that there is nothing left but empty shells of pasta. It's much more palatable, code-wise, than spaghetti (where the goodness is all over the plate and can't be stuffed in) or mud.
In the same way that there is duplication detection there should also be a metric available to judge whether a piece of code is not providing enough functionality or value. However, the way most code is written this isn't frequently a problem. A good example given of "about the right size" is where the algorithm is on paper and the code is in our head.
Wednesday, January 11, 2006
Dull Little Boxes
New Apple ad with about Intel Inside. From the unlikely URL of: http://www.apple.com/intel/ (which also includes a link to the keynote).
Tuesday, January 10, 2006
Northrop Grumman Killing Kowari?
Resignation from Kowari due to Northrop Grumman Letter "It is with sincere regret that I resigned as an administrator and developer of the Kowari Metastore."
"Northrop Grumman's position seems to be that they "purchased all rights associated with the Kowari software", a position not reconcilable with their continued release of the software under the Mozilla Public License, version 1.1."
So it would seem that Northrop, for whatever reason, dislikes Kowari's existence.
To get David and then Tucana (whose IP was eventually sold to Northrop) to start and continue with an open source version of TKS (the closed source version of the RDF database) was my little hobby horse (I fought fairly long and hard you could say). The idea was to get the Semantic Web bootstrapped and to create a value network around Semantic Web technology (I got this originally from "A critical look at object-orientation"). It is deeply disappointing for me to see it (the value network and Kowari) threatened in this way.
Like any OS project it is the ecosystem that is built around Kowari that makes it work. With so many contributions, just how much of the total code does Northrop actually own the copyright to? By accepting changes it isn't following the dual licencing model used by the owners of Berkley DB. To me Kowari was always using open source as a development and distribution model. Already, we have some contributors threatening to pull out their code (more at, "Disturbing News From Kowari").
Having administrators leave due to Northrop's actions is a good way to kill the project. Hopefully, Northrop's bullying will get a lot of interest - it's good to see, for instance, that this has been picked by aggregators like Topix's Aerospace and Defense.
The other posting David W linked to: Is Northrup Grumman Smushing Kowari?, "But I’d like to know what the status of Kowari is going to be, open source or not, because there are clients for whom Kowari might be the right choice, and we won’t bet on a dog that Northrup Grumman seems determined to publicly beat to death.
Relying on Kowari is now not prudent, given our obligation to do our best for our clients; but it’s also bad for Semantic Web uptake in the US federal government, and that’s something Northrup Grumman should think very carefully about."
Update: Paul and Andrae have also added their thoughts.
Update 2: David has published the letter from Northrop's lawyers. Danny posts, "...clearly if the current Kowari is MPL’d they can’t stop other developers working on the system."
"Northrop Grumman's position seems to be that they "purchased all rights associated with the Kowari software", a position not reconcilable with their continued release of the software under the Mozilla Public License, version 1.1."
So it would seem that Northrop, for whatever reason, dislikes Kowari's existence.
To get David and then Tucana (whose IP was eventually sold to Northrop) to start and continue with an open source version of TKS (the closed source version of the RDF database) was my little hobby horse (I fought fairly long and hard you could say). The idea was to get the Semantic Web bootstrapped and to create a value network around Semantic Web technology (I got this originally from "A critical look at object-orientation"). It is deeply disappointing for me to see it (the value network and Kowari) threatened in this way.
Like any OS project it is the ecosystem that is built around Kowari that makes it work. With so many contributions, just how much of the total code does Northrop actually own the copyright to? By accepting changes it isn't following the dual licencing model used by the owners of Berkley DB. To me Kowari was always using open source as a development and distribution model. Already, we have some contributors threatening to pull out their code (more at, "Disturbing News From Kowari").
Having administrators leave due to Northrop's actions is a good way to kill the project. Hopefully, Northrop's bullying will get a lot of interest - it's good to see, for instance, that this has been picked by aggregators like Topix's Aerospace and Defense.
The other posting David W linked to: Is Northrup Grumman Smushing Kowari?, "But I’d like to know what the status of Kowari is going to be, open source or not, because there are clients for whom Kowari might be the right choice, and we won’t bet on a dog that Northrup Grumman seems determined to publicly beat to death.
Relying on Kowari is now not prudent, given our obligation to do our best for our clients; but it’s also bad for Semantic Web uptake in the US federal government, and that’s something Northrup Grumman should think very carefully about."
Update: Paul and Andrae have also added their thoughts.
Update 2: David has published the letter from Northrop's lawyers. Danny posts, "...clearly if the current Kowari is MPL’d they can’t stop other developers working on the system."
Bookshelf of Bookmarks
- Slashdot is Going out of Style in 2006 "2006 will the year in which the once great Slashdot dies." More up-to-date and more stories over comments?
- Comparing LINQ and Its Contemporaries Comparison of Ruby on Rails, Hibernate and other languages/APIs and LINQ.
- Fire & Motion: The Trouble with Competing with Google How Google picks its fights in order to cause the most cost to its competitors.
- RDF For Rapid Development Of A Data Model How to use RDF to eventually get to a relational or other integration system (mentions Metis).
- RuleML "The aim is to build an expert system shell which implements native support for RuleML and incorporates the latest concepts in web-based and distributed processing." Via Sumatra Posted to Sourceforge.
- From the same site: Grand Challenge, Neural nets and rule engines How to distribute the RETE algorithm by processing locally and distributing indexes.
- A couple of links on the Ruby VM, Rite: Rite and YARV (Yet Another Ruby VM)
- Luke 6:41 "Our Leader hasn't caught Osama bin Laden, but he's doing a bang up job rounding up brown people."
- Modularity Rules "IBM’s investment changed the nature of the industry so completely that modularization is the name of the game."
- JAHAH Home AJAX is so 2005.
Sunday, January 08, 2006
Bug or Feature
ImplicitInterfaceImplementation "What would it mean to implement an implicit interface? Essentially it would tell the type system that the ValuedCustomer class implements all the methods declared in the public interface of Customer but does not take any of its implementation, that is its public method bodies, and non public methods or data. In other words we have interface-inheritance but not implementation-inheritance."
InterfaceImplementationPair "Using interfaces when you aren't going to have multiple implementations is extra effort to keep everything in sync (although good IDEs help). Furthermore it hides the cases where you actually do provide multiple implementations."
Interfaces "Somewhat related to this is one of Robert Martin's Principles of Object-Oriented Design: "The Interface Segregation Principle", which says "Clients should not be forced to depend on methods that they do not use. Interfaces belong to client, not to hierarchies."
In a dynamic system like Smalltalk, Ruby, or Python, we don't need explicit interfaces, so "The Interface Segregation Principle" doesn't seem to apply. In static systems like Java and C++, one could declare interfaces specific to each client, but it seems that few people do.
The principle seems to more about dependency-breaking than anything else. I'd be interested in hearing from people who have actually done this extensively."
Implicit Interfaces "Of course Martin's suggestion would make it easier to use third party code without wrapping it up in a shim of your own. Personally I think it's a good thing to get in the habit of segregating third party code and putting it behind a thin, object based firewall of your own devising. The shim/wrapper code allows you to blend the alien API into your standard ways of doing things and gives you a chance to refine the abstraction that the third party API provides. Chances are their abstractions are in terms of the problem that they solve whereas yours should be in terms of the problem that you are solving and might only use a small part of the API... What's more, if you have slipped in an interface you can test without the third party code and you can, potentially, replace the third party code with other code if required..."
This last point is very much in line with my own experience and the point in A RATIONAL DESIGN PROCESS: HOW AND WHY TO FAKE IT "Often we are encouraged, for economic reasons, to use software that was developed for some other
project...The resulting software may not be the ideal software for either project, i.e., not the software that we would develop based on its requirements alone, but it is good enough and will save effort."
InterfaceImplementationPair "Using interfaces when you aren't going to have multiple implementations is extra effort to keep everything in sync (although good IDEs help). Furthermore it hides the cases where you actually do provide multiple implementations."
Interfaces "Somewhat related to this is one of Robert Martin's Principles of Object-Oriented Design: "The Interface Segregation Principle", which says "Clients should not be forced to depend on methods that they do not use. Interfaces belong to client, not to hierarchies."
In a dynamic system like Smalltalk, Ruby, or Python, we don't need explicit interfaces, so "The Interface Segregation Principle" doesn't seem to apply. In static systems like Java and C++, one could declare interfaces specific to each client, but it seems that few people do.
The principle seems to more about dependency-breaking than anything else. I'd be interested in hearing from people who have actually done this extensively."
Implicit Interfaces "Of course Martin's suggestion would make it easier to use third party code without wrapping it up in a shim of your own. Personally I think it's a good thing to get in the habit of segregating third party code and putting it behind a thin, object based firewall of your own devising. The shim/wrapper code allows you to blend the alien API into your standard ways of doing things and gives you a chance to refine the abstraction that the third party API provides. Chances are their abstractions are in terms of the problem that they solve whereas yours should be in terms of the problem that you are solving and might only use a small part of the API... What's more, if you have slipped in an interface you can test without the third party code and you can, potentially, replace the third party code with other code if required..."
This last point is very much in line with my own experience and the point in A RATIONAL DESIGN PROCESS: HOW AND WHY TO FAKE IT "Often we are encouraged, for economic reasons, to use software that was developed for some other
project...The resulting software may not be the ideal software for either project, i.e., not the software that we would develop based on its requirements alone, but it is good enough and will save effort."
Friday, January 06, 2006
Mac XP
Sort of like the "Pooh and the Philosophers" only for XP and the Mac development team.
Project Metaphor: Creative Think "Find a central metaphor that's so good that everything aligns to it. Design meetings are no longer necessary, it designs itself. The metaphor should be crisp and fun."
Refactoring: -2000 Lines Of Code "He recently was working on optimizing Quickdraw's region calculation machinery, and had completely rewritten the region engine using a simpler, more general algorithm which, after some tweaking, made region operations almost six times faster. As a by-product, the rewrite also saved around 2,000 lines of code."
Simple Design: MacPaint Evolution "I was surprised a few days later when Bill told me that he decided to remove the character recognition feature from MacPaint. He was afraid that if he left it in, people would actually use it a lot, and MacPaint would be regarded as an inadequate word processor instead of a great drawing program. It was probably the right decision, although I didn't think so at the time. I was amazed that he was able to detach himself from all the effort that he put into creating the discarded feature; I know that I probably wouldn't have been able to do the same."
Constant Integration: Real Artists Ship "We starting doing release cycles that were only a few hours apart, re-releasing every time we fixed a significant problem."
Pair Programming: Real Artists Ship "At one point, around 2am on Sunday night, I stumbled across a bug in the clipboard code. I thought I knew what it might be, but I was so tired that I didn't want to deal with it. I tried to pretend that I didn't see the problem, but Steve Capps was watching my expression and knew there was something wrong. I also was too tired to sustain a pretense; he grilled me about the problem and then helped me craft a fix, since I was too tired to do it on my own."
Automated Testing: Monkey Lives "The Monkey was a small desk accessory that used the journaling hooks to feed random events to the current application, so the Macintosh seemed to be operated by an incredibly fast, somewhat angry monkey, banging away at the mouse and keyboard, generating clicks and drags at random positions with wild abandon. It had great potential as a testing tool, so Capps refined it to generate more semantically rich events, with a certain percentage of the events as menu commands, a certain percentage as window drags, etc."
Coding Standard: Hungarian "Bud decided that it would be too error prone to try to translate the Hungarian memory manager directly into assembly language. First, he made a pass through it to strip the type prefixes and restore the vowels to all the identifier names, so you could read the code without getting a headache, before adding lots of block comments to explain the purpose of various sub-components.
A few weeks later, when Bud came back to attend one of our first retreats, he brought with him a nicely coded, efficient assembly language version of the memory manager, complete with easy to read variable names, which immediately became a cornerstone of our rapidly evolving Macintosh operating system."
The importance of on-site customers: Round Rects Are Everywhere! "Steve suddenly got more intense. "Rectangles with rounded corners are everywhere! Just look around this room!". And sure enough, there were lots of them, like the whiteboard and some of the desks and tables. Then he pointed out the window. "And look outside, there's even more, practically everywhere you look!". He even persuaded Bill to take a quick walk around the block with him, pointing out every rectangle with rounded corners that he could find. ...Over the next few months, roundrects worked their way into various parts of the user interface, and soon became indispensable."
Project Metaphor: Creative Think "Find a central metaphor that's so good that everything aligns to it. Design meetings are no longer necessary, it designs itself. The metaphor should be crisp and fun."
Refactoring: -2000 Lines Of Code "He recently was working on optimizing Quickdraw's region calculation machinery, and had completely rewritten the region engine using a simpler, more general algorithm which, after some tweaking, made region operations almost six times faster. As a by-product, the rewrite also saved around 2,000 lines of code."
Simple Design: MacPaint Evolution "I was surprised a few days later when Bill told me that he decided to remove the character recognition feature from MacPaint. He was afraid that if he left it in, people would actually use it a lot, and MacPaint would be regarded as an inadequate word processor instead of a great drawing program. It was probably the right decision, although I didn't think so at the time. I was amazed that he was able to detach himself from all the effort that he put into creating the discarded feature; I know that I probably wouldn't have been able to do the same."
Constant Integration: Real Artists Ship "We starting doing release cycles that were only a few hours apart, re-releasing every time we fixed a significant problem."
Pair Programming: Real Artists Ship "At one point, around 2am on Sunday night, I stumbled across a bug in the clipboard code. I thought I knew what it might be, but I was so tired that I didn't want to deal with it. I tried to pretend that I didn't see the problem, but Steve Capps was watching my expression and knew there was something wrong. I also was too tired to sustain a pretense; he grilled me about the problem and then helped me craft a fix, since I was too tired to do it on my own."
Automated Testing: Monkey Lives "The Monkey was a small desk accessory that used the journaling hooks to feed random events to the current application, so the Macintosh seemed to be operated by an incredibly fast, somewhat angry monkey, banging away at the mouse and keyboard, generating clicks and drags at random positions with wild abandon. It had great potential as a testing tool, so Capps refined it to generate more semantically rich events, with a certain percentage of the events as menu commands, a certain percentage as window drags, etc."
Coding Standard: Hungarian "Bud decided that it would be too error prone to try to translate the Hungarian memory manager directly into assembly language. First, he made a pass through it to strip the type prefixes and restore the vowels to all the identifier names, so you could read the code without getting a headache, before adding lots of block comments to explain the purpose of various sub-components.
A few weeks later, when Bud came back to attend one of our first retreats, he brought with him a nicely coded, efficient assembly language version of the memory manager, complete with easy to read variable names, which immediately became a cornerstone of our rapidly evolving Macintosh operating system."
The importance of on-site customers: Round Rects Are Everywhere! "Steve suddenly got more intense. "Rectangles with rounded corners are everywhere! Just look around this room!". And sure enough, there were lots of them, like the whiteboard and some of the desks and tables. Then he pointed out the window. "And look outside, there's even more, practically everywhere you look!". He even persuaded Bill to take a quick walk around the block with him, pointing out every rectangle with rounded corners that he could find. ...Over the next few months, roundrects worked their way into various parts of the user interface, and soon became indispensable."
Wednesday, January 04, 2006
New Newsy News
I should've done this a little while ago, a link to The News before The News "Tech opinion from James Webster". Perhaps one of the few people to constantly surprise me with stuff that I hadn't read about before.
Includes links to my blog including: "All Yuor Google Base Are Belong To Us!" and "DabbleDb - Naked Objects for Web 2.0".
Includes links to my blog including: "All Yuor Google Base Are Belong To Us!" and "DabbleDb - Naked Objects for Web 2.0".
Tuesday, January 03, 2006
Rocketman
First jet powered Birdman flight "After checking the altimeter several times, it was apparent that there was no appreciable loss in altitude for this period of time. Visa next changed his angle of attack by redirected the thrust and changing his body position to attain vertical climb...The jump has proven empirically that level human flight is possible and sustainable using the combination of jet engines and a bird-man suit. The strength required to control level flight was relatively easy, and controlling the direction of flight feels surprisingly natural. The duration of flight is simply a factor of the consumption of fuel of the engine(s) powering the flight."
Video here.
Video here.
Saturday, December 31, 2005
Smalltalk meets the Semantic Web
Smalltalk:::OWL-Project "OWL has emerged from the AI/semantic community and tends to be in the open-source community which appears to be a direction for Smalltalk (e.g. Smalltalk Solutions at Linux World) Much of the work to date has been implemented in Python and Ruby which, from a language perspective, is very close to Smalltalk. However, those languages become less appealing if you have ever worked in the IDE's supporting those languages. OWL can provide the Smalltalk community with a "market" that is a good fit for the features of the ST language and supporting IDE's."
"Agilense provides a product named EA WebModeler...is an implementation of the Adaptive Object Model pattern..."
A good summary of AOM: "We call these systems "Adaptive Object-Models", because the users' object model is interpreted at runtime and can be changed with immediate (but controlled) effects on the system interpreting it. The real power in Adaptive Object-Models is that the definition of a domain model and rules for its integrity can be configured by domain experts external to the execution of the program."...If you got just one reason why EJB is flawed, this has got to be the one, you can't build systems for Enterprises at the same time take away control form the business stakeholders."
"Agilense provides a product named EA WebModeler...is an implementation of the Adaptive Object Model pattern..."
A good summary of AOM: "We call these systems "Adaptive Object-Models", because the users' object model is interpreted at runtime and can be changed with immediate (but controlled) effects on the system interpreting it. The real power in Adaptive Object-Models is that the definition of a domain model and rules for its integrity can be configured by domain experts external to the execution of the program."...If you got just one reason why EJB is flawed, this has got to be the one, you can't build systems for Enterprises at the same time take away control form the business stakeholders."
RETE Rebuked
A recent blog I started reading after criticizing SPARQL. This time its criticizing an entry Paul made about the scalability of the RETE algorithm: "I am a bit confused by the statement that RETE does not scale. This is contrary to a mountain of papers by researchers and developers around the world. From the paragraph, a couple of things come to mind. "(loading indexes) does not need to be done often" tells me the author doesn't understand the purpose and goal of RETE. RETE was designed to solve machine learning problems where data changes rapidly and reasoning is a continuous process. What the author wants is something closer to BitMap indexes used in OLAP products."
A lot of the points raised, like RETE being for changing data, are mentioned in a previous post under Meeting. Drools was chosen as a starting to point to see what kind of system needed to be developed in Kowari.
Also mentioned, bitmap indexing: "Given Tucana is indexing everything, they might as well adapt Bitmap indexing and get better than linear performance. The problem described by the blog is a well understood problem in the OLAP world."
An previous entry, "Relational theory, RETE and Derby" points to some interesting articles about bitmap indexes (available in Oracle 9) and high scalability requirements: "In a large financial institution like a mutual fund company, they may have 1-20 million customers. If each customer has an average of 20-30 positions (aka specific holding of an equity) that means the potential dataset for firm wide compliance rule could involve 20million+ rows. Doing this within 2-5 seconds is rather hard, so it requires using lots of different techniques. In the extreme cases, a company might have 20 million accounts, which means the potential dataset is 600 million rows."
From the OTN article: "B-tree indexes are usually used when columns are unique or near unique; bitmap indexes should be used, or at least considered, in all other cases. While you would not generally use a b-tree index when retrieving 40 percent of the rows of a table, a bitmap index is often still faster than doing a full table scan. This is seemingly in violation of the 80/20 rule, which is to generally use an index when retrieving 20 percent or less of the rows and do a full table scan when retrieving more. Bitmap indexes are smaller and work differently from the 80/20 rule. You can effectively use bitmap indexes even when retrieving large percentages (20 to 80 percent) of a table. Bitmaps can also be used to retrieve conditions based on nulls (since nulls are also indexed) and for "not equal" conditions."
It would appear that this would be suitable for predicate indexation but not generally as both subjects and objects are near unique.
A lot of the points raised, like RETE being for changing data, are mentioned in a previous post under Meeting. Drools was chosen as a starting to point to see what kind of system needed to be developed in Kowari.
Also mentioned, bitmap indexing: "Given Tucana is indexing everything, they might as well adapt Bitmap indexing and get better than linear performance. The problem described by the blog is a well understood problem in the OLAP world."
An previous entry, "Relational theory, RETE and Derby" points to some interesting articles about bitmap indexes (available in Oracle 9) and high scalability requirements: "In a large financial institution like a mutual fund company, they may have 1-20 million customers. If each customer has an average of 20-30 positions (aka specific holding of an equity) that means the potential dataset for firm wide compliance rule could involve 20million+ rows. Doing this within 2-5 seconds is rather hard, so it requires using lots of different techniques. In the extreme cases, a company might have 20 million accounts, which means the potential dataset is 600 million rows."
From the OTN article: "B-tree indexes are usually used when columns are unique or near unique; bitmap indexes should be used, or at least considered, in all other cases. While you would not generally use a b-tree index when retrieving 40 percent of the rows of a table, a bitmap index is often still faster than doing a full table scan. This is seemingly in violation of the 80/20 rule, which is to generally use an index when retrieving 20 percent or less of the rows and do a full table scan when retrieving more. Bitmap indexes are smaller and work differently from the 80/20 rule. You can effectively use bitmap indexes even when retrieving large percentages (20 to 80 percent) of a table. Bitmaps can also be used to retrieve conditions based on nulls (since nulls are also indexed) and for "not equal" conditions."
It would appear that this would be suitable for predicate indexation but not generally as both subjects and objects are near unique.
Thursday, December 29, 2005
Good API Design
Java API Design Guidelines "If your API is worth anything, it will evolve over time...decide what sort of compatibility you will guarantee between revisions...What should the design goals of your API be?...absolutely correct...easy to use...easy to learn...fast enough...small enough...it's much easier to put things in than to take them out."
This seems all well and good.
"Interfaces can be implemented by anybody. Suppose String were an interface. Then you could never be sure that a String you got from somewhere obeyed the semantics you expect: it is immutable; its hashCode() is computed in a certain way; its length is never negative; and so on."
Adding interfaces on top of String like CharSequence was a good thing. It meant that you could process a String or a StringBuffer the same as well as consistently treat something that may have been in memory or on-disk (including being memory mapped via NIO). The way to ensure that String implementations do follow the correct semantics is test driving the interfaces and this requires being able to create a Mock object of that interface - pretty much sealing the deal as far as interfaces are concerned.
Actually, the whole section on interfaces is pretty much a wash as this can be solved by test driving 3 of the 4 points raised. As far as using an abstract class to help with the evolution of the API, you can have both - an interface and a default abstract class for some base implementation.
Exceptions get a going over too, "Use a checked exception "if the exceptional condition cannot be prevented by proper use of the API and the programmer using the API can take some useful action once confronted with the exception." In practice this usually means that a checked exception reflects a problem in interaction with the outside world, such as the network, filesystem, or windowing system."
Related: Evolution Not Creation, Reminder About Incremental and Test Driven Development and 10 Minute Commits for Better Code.
This seems all well and good.
"Interfaces can be implemented by anybody. Suppose String were an interface. Then you could never be sure that a String you got from somewhere obeyed the semantics you expect: it is immutable; its hashCode() is computed in a certain way; its length is never negative; and so on."
Adding interfaces on top of String like CharSequence was a good thing. It meant that you could process a String or a StringBuffer the same as well as consistently treat something that may have been in memory or on-disk (including being memory mapped via NIO). The way to ensure that String implementations do follow the correct semantics is test driving the interfaces and this requires being able to create a Mock object of that interface - pretty much sealing the deal as far as interfaces are concerned.
Actually, the whole section on interfaces is pretty much a wash as this can be solved by test driving 3 of the 4 points raised. As far as using an abstract class to help with the evolution of the API, you can have both - an interface and a default abstract class for some base implementation.
Exceptions get a going over too, "Use a checked exception "if the exceptional condition cannot be prevented by proper use of the API and the programmer using the API can take some useful action once confronted with the exception." In practice this usually means that a checked exception reflects a problem in interaction with the outside world, such as the network, filesystem, or windowing system."
Related: Evolution Not Creation, Reminder About Incremental and Test Driven Development and 10 Minute Commits for Better Code.
5 Minutes with Monad
I've recently spent a bit of time trying to solve some problems using Microsoft's new Monad shell. It's interesting that the creature for the O'Reilly book is the common toad. It was a different experience, although the pain is eased a little as the default installation comes with all the new commands (cmdlets) mapped to Unix ones (like ls and ps).
The default security setting prevents remotely signed objects from being executed and there seems to be no way to turn it on. The documentation is missing. To make it usable it's:
Going through the tutorials it did show itself to be kind of cool. For example, being able to select the top 10 processes based on VirtualMemorySize:
One of the problems was trying to do line by line processing. There was promises of pipelining via XML streams but according to "Replace lines in a text file?" (the first hit on Google) Monad doesn't support it. The lack of streaming appears to be a crucial omission in a toolset designed for system administrators - although it might not be fatal as log files and the like don't usually come close to the available memory of modern systems.
It does support accessing the .NET APIs which provides a loophole. For example, to read a file line by line and replace "xxx" with "yyy":
It was all for nothing, as I later found out that it didn't support Windows 2000 and it needed to be deployed on that - it is supported by Windows XP, 2003 and Vista. Back to Windows Script Host (maybe using Ruby) I guess.
The default security setting prevents remotely signed objects from being executed and there seems to be no way to turn it on. The documentation is missing. To make it usable it's:
set-property
`HKLM:\SOFTWARE\Microsoft\Msh\Microsoft.Management.Automation.msh`
-property ExecutionPolicy -value RemoteSigned
Going through the tutorials it did show itself to be kind of cool. For example, being able to select the top 10 processes based on VirtualMemorySize:
get-process | sort-object VirtualMemorySize | select-object -last 10You can whack on a "convert-HTML" or an "export-csv" to produce the result in a format you want or connect to Excel or SQL Server to retrieve data. A lot has been made of its native XML support and how it passes around strongly typed objects rather than just Unix's streams.
One of the problems was trying to do line by line processing. There was promises of pipelining via XML streams but according to "Replace lines in a text file?" (the first hit on Google) Monad doesn't support it. The lack of streaming appears to be a crucial omission in a toolset designed for system administrators - although it might not be fatal as log files and the like don't usually come close to the available memory of modern systems.
It does support accessing the .NET APIs which provides a loophole. For example, to read a file line by line and replace "xxx" with "yyy":
$f = [System.IO.File]::OpenText("c:\file.txt")
while($line = $f.ReadLine())
{
$line -replace "xxx","yyy"
}It was all for nothing, as I later found out that it didn't support Windows 2000 and it needed to be deployed on that - it is supported by Windows XP, 2003 and Vista. Back to Windows Script Host (maybe using Ruby) I guess.
Sunday, December 25, 2005
Merry Bag Of Links
Political:
Agile:
Programming:
General Technical:
- Support Creative Commons "We are down to the last $100,000, and really need your support — both for the very cool projects we’re launching (see, e.g., the license interoperability project, discussed recently in Technology Review, and the two new projects announced this week), and for the very uncool pressure we’re under from IRS regulations to demonstrate “public support” as a condition for keeping our (absolutely essential as in we can’t live with out it) tax exempt status." Via We've got 10 days, and we need $100,000. Please help
- Passion of the Spaghetti Monster and Intelligent Design
- Top 12 media myths and falsehoods on the Bush administration's spying scandal "...the Bush administration and its conservative allies in the media have defended the secret spying operation with false and misleading claims that have subsequently been reported without challenge across the media."
- The Curious Section 126 of the Patriot Act "Congress is seeking assurances that "the privacy and due process rights of individuals" is protected in the course of the government using massive databases of non-publicly available data; both proprietary databases and its own compiled intelligence and law enforcement databases to "search" for terrorists and terrorist connections."
Agile:
- Client vs. Developer Wars "This, to me, is another indictment of dysfunctional specifications. I learned long ago that clients won't listen to what you say, and they certainly won't read what you write. You're much better off putting that wasted effort into a working model and setting it in front of the client. Let them play with it for a while. Refine the working model based on that feedback, then keep turning the crank on this cycle until you run out of resources."
- Continuous Testing - in spirals "I want tests to run ‘inside out’ Imagine a spiral, with the unit test for the bit of code you’re currently editing to be the focal point. Ideally, I’d want a test to run first for the method I changed last, then for the whole class, then for the suite the class is in, then further out to other dependencies. Tests run outward only when green bars are encountered. If there is a red bar somewhere, the spiraling stops, so we can examine the failure, fix it, and see again from the inside which tests run."
- Essential Advice for Agile Coaches
Programming:
- Ruby Off the Rails "Ruby's syntax is quite different from that of the Java language, but it's amazingly easy to pick up. Moreover, some things are just plain easier to do in Ruby than they are in the Java language."
- Seven Habits of Highly Effective Programmers "Sorry, there's no shortcut - you have to learn and practice and make some mistakes."
- Networking Libraries using NIO: EJOE, Coconut AIO and MINA. MINA's features include: "unit testability using mock objects".
- ONJava: 2005 Year in Review "Java is still by far the most widely used programming language, if book sales are any indication, about 2x C#, 2.5x PHP, 4x Perl, and 9x Ruby/Python."
- Automate acceptance tests with Selenium.
General Technical:
- What is Songbird? "Songbird is built atop the Mozilla Foundation's XULRunner platform also used by the Firefox browser, the Thunderbird email client and other desktop applications." The user preview has been delayed
- Mac Mini - Big Ideas
- Lack of focus and death march at Google
- Mac IE's Death: A Case for Microsoft Disbanding or Transfering the Windows IE Team ""Then why on earth did we pursue IE in the first place? Just so that the DOJ would sue us?""
- The Bubble Cycle is Replacing the Business Cycle Housing, bond, etc bubbles.
Saturday, December 24, 2005
Know When to Hold Them, Know When to Fold Them
So what is an RDF merge and when should you apply it in SPARQL?
Very succinctly, in "The Semantics of SPARQL" it says: "The RDF merge
<G1...Gn> of a sequence of graphs <G1...Gn> (i.e., a dataset) is the ordered merge union of the graphs, where repeated bnodes are substituted with fresh ones, by keeping the names of the bnodes coming first in the sequence order."
In "SPARQL Query Language for RDF" it gives a simple example:
Graph 1:
_:a foaf:name "Bob" .
_:a foaf:mbox .
Graph 2:
_:a foaf:name "Alice" .
_:a foaf:mbox .
The result of the merge, upon which queries are made:
_:x foaf:name "Bob" .
_:x foaf:mbox .
_:y foaf:name "Alice" .
_:y foaf:mbox .
Section 9 details querying multiple graphs in SPARQL, including a new dataset where the default graph is a merge of the graphs in the FROM clause.
In summary, when SPARQL operations are performed across graphs you get new blank nodes which prevents, for example, being able to JOIN across graphs using them.
What is generally required by RDF applications is something like smushing. For example, an "...RDF spider (often known as a "scutter") can gather up FOAF files and "smush" them together into a single model that unifies the individual pieces of information into a network." (from "A Semantic Web Shoebox - Annotating Photos with RSS and RDF").
To actually achieve smushing, Leo has an example algorithm or it might be appropriate to adapt RDF graph isomorphism algorithms.
Very succinctly, in "The Semantics of SPARQL" it says: "The RDF merge
In "SPARQL Query Language for RDF" it gives a simple example:
Graph 1:
_:a foaf:name "Bob" .
_:a foaf:mbox
Graph 2:
_:a foaf:name "Alice" .
_:a foaf:mbox
The result of the merge, upon which queries are made:
_:x foaf:name "Bob" .
_:x foaf:mbox
_:y foaf:name "Alice" .
_:y foaf:mbox
Section 9 details querying multiple graphs in SPARQL, including a new dataset where the default graph is a merge of the graphs in the FROM clause.
In summary, when SPARQL operations are performed across graphs you get new blank nodes which prevents, for example, being able to JOIN across graphs using them.
What is generally required by RDF applications is something like smushing. For example, an "...RDF spider (often known as a "scutter") can gather up FOAF files and "smush" them together into a single model that unifies the individual pieces of information into a network." (from "A Semantic Web Shoebox - Annotating Photos with RSS and RDF").
To actually achieve smushing, Leo has an example algorithm or it might be appropriate to adapt RDF graph isomorphism algorithms.
Friday, December 23, 2005
More on Disjunction
I think this is just going to happen time and time again, "RDF non-sense": "The W3C Working Group members argue SPARQL is a mixed mode language that does support OR, though they are calling it an "optional union". Frankly, I see no point in renaming something people understand to mean one thing. It gives me the impression the W3C want to be thought police and enforce a certain way of thinking.
On the practical side, many analysts happen to like OR disjunctions and would complain loudly. Of course, there are plenty of cases where users abuse the power and write deeply nested disjunctions. That is not a valid reason in my mind to avoid disjunction. It saves the user time and allows them to write simpler rules using disjunction. The W3C seems love RDF and wants the world to love it. Unfortunately, the current specification is a complete piece of junk. I hope RDF dies a quick and public death."
I declare...backward chaining suits me fine! "My two favorite declarative tools right now are Pellet and Prova, both of which are open source java SemWeb tools that are highly compatible with Jena, which recently got a bump to 2.3 with fairly complete SPARQL support. Pellet is an implementation of OWL-DL and some related description logic facilities by the Mindswap guys in Maryland, who have absorbed some of the l33t Kowari/Tucana guys, too (Tucana was recently picked up by Northrop, BTW)."
"Prova is a prolog-variant built on top of Mandarax. It is a very effective and fun medium for scripting of high-level relationships and operations. The integration of prolog unification, java types, java methods, and java exceptions is done very nicely, and yields fine code economy. There are some rough edges in the docs, but we are helping to get these worked out in the pretty soon."
I guess that means David W is l33t! :-)
On the practical side, many analysts happen to like OR disjunctions and would complain loudly. Of course, there are plenty of cases where users abuse the power and write deeply nested disjunctions. That is not a valid reason in my mind to avoid disjunction. It saves the user time and allows them to write simpler rules using disjunction. The W3C seems love RDF and wants the world to love it. Unfortunately, the current specification is a complete piece of junk. I hope RDF dies a quick and public death."
I declare...backward chaining suits me fine! "My two favorite declarative tools right now are Pellet and Prova, both of which are open source java SemWeb tools that are highly compatible with Jena, which recently got a bump to 2.3 with fairly complete SPARQL support. Pellet is an implementation of OWL-DL and some related description logic facilities by the Mindswap guys in Maryland, who have absorbed some of the l33t Kowari/Tucana guys, too (Tucana was recently picked up by Northrop, BTW)."
"Prova is a prolog-variant built on top of Mandarax. It is a very effective and fun medium for scripting of high-level relationships and operations. The integration of prolog unification, java types, java methods, and java exceptions is done very nicely, and yields fine code economy. There are some rough edges in the docs, but we are helping to get these worked out in the pretty soon."
I guess that means David W is l33t! :-)
I Hope Not
Ruby is to Perl what C++ was to C. He qualified this by saying, "Ruby improves and simplifies the Perl language" which isn't really what I saw C++ doing at all.
Quote: "These are the folks that assert that Java's verbosity is "just finger typing that Eclipse/IntelliJ will do for me," and it doesn't matter if the resulting code has 20 times the visual bulk of a simpler approach. One of the basic tenets of the Python language has been that code should be simple and clear to express and to read, and Ruby has followed this idea, although not as far as Python has because of the inherited Perlisms. But for someone who has invested Herculean effort to use EJBs just to baby-sit a database, Rails must seem like the essence of simplicity. The understandable reaction for such a person is that everything they did in Java was a waste of time, and that Ruby is the one true path."
Related: Rocking With Ruby and One way Java is better than Ruby
Quote: "These are the folks that assert that Java's verbosity is "just finger typing that Eclipse/IntelliJ will do for me," and it doesn't matter if the resulting code has 20 times the visual bulk of a simpler approach. One of the basic tenets of the Python language has been that code should be simple and clear to express and to read, and Ruby has followed this idea, although not as far as Python has because of the inherited Perlisms. But for someone who has invested Herculean effort to use EJBs just to baby-sit a database, Rails must seem like the essence of simplicity. The understandable reaction for such a person is that everything they did in Java was a waste of time, and that Ruby is the one true path."
Related: Rocking With Ruby and One way Java is better than Ruby
Wednesday, December 21, 2005
A Better PageRank
This paper gives an example of some of the flaws with Google's PageRank algorithm and they suggest they have an algorithm that fixes it. Something is Wrong with Google’s Mathematical Model "In their original paper : ”The PageRank Citation Ranking: Bringing Order to the Web” [1], Page et al. suggest a new ranking algorithm named - PageRank. It is shown there that implementing the new algorithm boils down to solving a huge eigenvalues problem Ax = x (1) where A is a matrix which represents a graph related to the web. It is claimed that in order for the model to work properly, the graph should be strongly connected. In general, the graph is not strongly connected and we have ’sink’ set of pages."
"We have developed a new algorithm which can be considered as a modification of the original PageRank algorithm. The modified algorithm is stable and gives a correct ranking vector. The mathematical complexity of the suggested algorithm is the same as the complexity of the original one."
In Google's Librarian Center they've published as similar explanation of PageRank called "How does Google collect and rank results?".
"We have developed a new algorithm which can be considered as a modification of the original PageRank algorithm. The modified algorithm is stable and gives a correct ranking vector. The mathematical complexity of the suggested algorithm is the same as the complexity of the original one."
In Google's Librarian Center they've published as similar explanation of PageRank called "How does Google collect and rank results?".
LiveConnect, Lives as LAJAX
I ignored this when I first heard about it, as noted here: ""I was disappointed when [Bray] said that he was going to make a product announcement, and I was unenthusiastic when the announcement turned out to be about a Sun version of Derby," Leung wrote.".
Derby Demo hits a nerve "I think we hit a nerve with this demo. I think many of us within the Derby community recognized the potential for Derby within a web browser environment, but it's wonderful, great, fantastic to see how the community is "getting" it and running with it."
Derby ApacheCon demo and ApacheCon 2005: Ok, Tim, I'm not jaded anymore.
Derby Demo hits a nerve "I think we hit a nerve with this demo. I think many of us within the Derby community recognized the potential for Derby within a web browser environment, but it's wonderful, great, fantastic to see how the community is "getting" it and running with it."
Derby ApacheCon demo and ApacheCon 2005: Ok, Tim, I'm not jaded anymore.
Monday, December 19, 2005
Minimum Union
The key operation to provide outer joins that are associative and communitive, as noted in Outer Joins Aren't Primitive, is called "minimum union". Minimum union (I've also seen "outer union") pads with nulls the tuples of two schemas and then unions them (without duplicate removal).
The original thesis "Algebraic Optimization of Outerjoin Queries" gives a variant of relational algebra that allows tuples defined by different sets of attributes (schemes) rather than padding with nulls. This actually removes the requirement for nulls. It also includes presenting the data in a nested relational form, where instead of having one tuple in a parent-child relationship, children are a set to a parent. This is just like the way Kowari presents its results (Figure 3.2 in the paper).
The original thesis "Algebraic Optimization of Outerjoin Queries" gives a variant of relational algebra that allows tuples defined by different sets of attributes (schemes) rather than padding with nulls. This actually removes the requirement for nulls. It also includes presenting the data in a nested relational form, where instead of having one tuple in a parent-child relationship, children are a set to a parent. This is just like the way Kowari presents its results (Figure 3.2 in the paper).
Sunday, December 18, 2005
A Link in Time
- Getting more than what you pay for with open source databases. The tale of using a free database, using a commercial one and then putting the free one back. A quote: "There are a lot of companies out there paying dearly for commercial databases (and operating systems for that matter). As far as I'm concerned they might as well be flushing that money down the toilet. Actually, they might be better off. We certainly would have been."
- For the metrically minded, "Two Motivational Metrics for Agile Teams". So as long as you convince people that "Time to Obstacle Removal" and "Obstacles Removed per Iteration" are good metrics.
- The, "I'm sick of the comparison between building and software" rant.
- Martin Fowler has recently updated his "The New Methodology" with half a dozen or so flavours of agile, which one are you?
- Don't Click It for those people with zero button mice. The first example of a gesture to replace a click just seems ridiculously hard.
- Examples of using Chickenfoot for Firefox: Rewrite the Web
The table sort, how to get a book from the library, simple examples and timely. - Another REST framework - Restlet.
- Getting stuff done - the quick version
- How not to get a PhD
Event Horizon
- Semantic Web, Here We Come "The “Structured Blogging Initiative” is an attempt to jump-start the “semantic web,” the idea of giving deeper meaning to the Internet advocated by World Wide Web creator Tim Berners-Lee. By incorporating descriptive information into the code of web pages, laypeople will be able to designate their content as a movie review, an event posting, or an item available for sale." StructureBlogging initiative and other entries: "Structured blogging initiative taking off", "More StructuredBlogging feedback" and Structured Blogging is a thing you do -- not a format.
- Bill de hÓra discusses RDF and database schemas: "Using RDF storage provides flexibility at the domain level. Altering tables isn't needed because RDF, being a graph based, is naturally additive...My (somewhat anecdotal) experience with RDF is that datasets in the order of 106 and greater aren't uncommon and that you should budget for an order of magnitude increase in terms of the number of rows required for the domain storage compared to an entity relational approach...It's an interesting question whether using RDBMSes to store RDF counts as some form of abuse, or bad engineering."
Saturday, December 17, 2005
Ion Inside
First Mass Producible Quantum Computer Chip "Using the same semiconductor fabrication technology that is used in everyday computer chips, researchers were able to trap a single atom within an integrated semiconductor chip and control it using electrical signals, said Christopher Monroe, U-M physics professor and the principal investigator and co-author of the paper, "Ion Trap in a Semiconductor Chip." The paper appeared in the Dec. 11 issue of Nature Physics."
Thursday, December 15, 2005
Marks
The Man Who Wasn't There: problems of missing or partially missing data in geoscience databases "In the literature discussions between Codd and Date on the propriety or otherwise of NULLs in relational databases, there seems to have been some confusion on both sides, on one very important question. That is the distinction between database representation and function evaluation. NULLs are one approach to the problem of handling missing data within the database."
So in this respect RDF is great - you don't have to come up with a value or values to represent missing data. You only have to worry about function evaluation.
"In fact the Codd 'mark' solution does not in itself require, as unfortunately implied by Codd himself, and vigorously attacked by Date, the use of 3- or 4-valued logic, and therefore cannot be dismissed so easily. Relational database theory is based on first-order predicate logic, which uses two truth values TRUE and FALSE. If there is no value for a data item, then the logical statement corresponding to the tuple containing that item can simply omit any mention of that particular column. If the value of this data item is required in an operation, then there is only one truth value which can be returned: FALSE. This applies to database set operations such as JOINs and also to numerical operations such as totals and averages where the absence of any required data value prevents the computation from being carried out. If a total or average is required in such a situation, then the problem can be circumvented only by first selecting non-absent data. This is the correct treatment, to ensure that statistics are computed on a valid data set."
Another example of marks, tuple marks.
SH writes in about incomplete data in observational science databases, the open world assumption, 3VL and NULL, McGoveran responds saying: "In a scientific database such as the type to which you allude, a reasonable interpretation of True and False under CWA is "valid by experiment and consistent with hypotheses" and "not validated by experiment or inconsistent with hypotheses". If you give this differentiation up with CWA and nulls, you've given up scientific reasoning and the scientific method."
In Kowari, if you have the following triples: _b1, <urn:sno>, "S1"; _b2, <urn:sno>, "S2"; _b3, <urn:pno> "P1".
And performed the following query:
It returns: _b1, _b2, null (really unconstrained).
However, if you select $s2 instead it returns: null, b3. Using the above idea, it would return _b1, _b2 for the first query and _b3 for the second.
I'm not sure I really like this solution, preserving unknown seems to make more sense as demonstrated in How FirstSQL Solves the EXISTS and Other Problems.
So in this respect RDF is great - you don't have to come up with a value or values to represent missing data. You only have to worry about function evaluation.
"In fact the Codd 'mark' solution does not in itself require, as unfortunately implied by Codd himself, and vigorously attacked by Date, the use of 3- or 4-valued logic, and therefore cannot be dismissed so easily. Relational database theory is based on first-order predicate logic, which uses two truth values TRUE and FALSE. If there is no value for a data item, then the logical statement corresponding to the tuple containing that item can simply omit any mention of that particular column. If the value of this data item is required in an operation, then there is only one truth value which can be returned: FALSE. This applies to database set operations such as JOINs and also to numerical operations such as totals and averages where the absence of any required data value prevents the computation from being carried out. If a total or average is required in such a situation, then the problem can be circumvented only by first selecting non-absent data. This is the correct treatment, to ensure that statistics are computed on a valid data set."
Another example of marks, tuple marks.
SH writes in about incomplete data in observational science databases, the open world assumption, 3VL and NULL, McGoveran responds saying: "In a scientific database such as the type to which you allude, a reasonable interpretation of True and False under CWA is "valid by experiment and consistent with hypotheses" and "not validated by experiment or inconsistent with hypotheses". If you give this differentiation up with CWA and nulls, you've given up scientific reasoning and the scientific method."
In Kowari, if you have the following triples: _b1, <urn:sno>, "S1"; _b2, <urn:sno>, "S2"; _b3, <urn:pno> "P1".
And performed the following query:
select $s1
...
where $s1 <urn:sno> $o1 or $s2 <urn:pno> $o2;
It returns: _b1, _b2, null (really unconstrained).
However, if you select $s2 instead it returns: null, b3. Using the above idea, it would return _b1, _b2 for the first query and _b3 for the second.
I'm not sure I really like this solution, preserving unknown seems to make more sense as demonstrated in How FirstSQL Solves the EXISTS and Other Problems.
Outer Joins aren't Primitive
Optional data in SPARQL seems to be equivalent to left outer join in SQL. As it turns out, outer joins can be composed of disjunctions. This is similar to the original MAYBE function suggested to be added to Kowari (although that suggestions is quite a deal simpler). The below paper outlines algorithms to do outer queries more efficiently. They require computing the anti-join of certain relations - an antij-oin being the set difference between two tables (or MINUS operation). Here is a good explanation of semi-joins and anti-joins.
Outerjoins as Disjunctions "The outerjoin operator is currently available in the query language of several major DBMSs, and it is included in the proposed SQL2 standard draft. However, “associativity problems” of the operator have been pointed out since its introduction. In this paper we propose a shift in the intuition behind outerjoin: Instead of computing the join while also preserving its arguments, outerjoin delivers tuples that come either from the join or from the arguments. Queries with joins and outerjoins deliver tuples that come from one out of several joins, where a single relation is a trivial join. An advantage of this view is that, in contrast to preservation, disjunction is commutative and associative, which is a significant property for intuition, formalisms, and generation of execution plans.Based on a disjunctive normal form, we show that some data merging queries cannot be evaluated by means of binary outerjoins, and give alternative procedures to evaluate those queries. We also explore several evaluation strategies for outerjoin queries, including the use of semijoin programs to reduce base relations."
Also related, Outer Join in Edutella where each part of the outer query is done individually.
Outerjoins as Disjunctions "The outerjoin operator is currently available in the query language of several major DBMSs, and it is included in the proposed SQL2 standard draft. However, “associativity problems” of the operator have been pointed out since its introduction. In this paper we propose a shift in the intuition behind outerjoin: Instead of computing the join while also preserving its arguments, outerjoin delivers tuples that come either from the join or from the arguments. Queries with joins and outerjoins deliver tuples that come from one out of several joins, where a single relation is a trivial join. An advantage of this view is that, in contrast to preservation, disjunction is commutative and associative, which is a significant property for intuition, formalisms, and generation of execution plans.Based on a disjunctive normal form, we show that some data merging queries cannot be evaluated by means of binary outerjoins, and give alternative procedures to evaluate those queries. We also explore several evaluation strategies for outerjoin queries, including the use of semijoin programs to reduce base relations."
Also related, Outer Join in Edutella where each part of the outer query is done individually.
Tuesday, December 13, 2005
A Really Interactive Query Language
SQLBuilder "SQLBuilder uses clever overriding of operators to make Python expressions build SQL expressions -- so long as you start with a Magic Object that knows how to fake it."
An example:
Via, SQL API "I'd rather SQLObject be built on some ORM-neutral layer, where you can move down to that layer when SQLObject doesn't fit your problem; as opposed to now, where you kind of have to work around SQLObject."
This is almost exactly like something I was thinking about, to prevent semantically incorrect SQL queries. Add an AJAX interface on this and it would be cool and useful.
An example:
>>> from SQLBuilder import *
>>> person = table.person
# person is now equivalent to the Person.q object from the SQLObject
# documentation
>>> person
person
>>> person.first_name
person.first_name
>>> person.first_name == 'John'
person.first_name = 'John'
Via, SQL API "I'd rather SQLObject be built on some ORM-neutral layer, where you can move down to that layer when SQLObject doesn't fit your problem; as opposed to now, where you kind of have to work around SQLObject."
This is almost exactly like something I was thinking about, to prevent semantically incorrect SQL queries. Add an AJAX interface on this and it would be cool and useful.
DRY and Embedded Program Code
What if you could define a user interface and surf it via a telephone or browser and that the data and state from one to the other was able to be shared across multiple users?
Beyond interactive voice response "English’s hope, he tells Inskeep in the interview, is that companies will admit how infuriating their systems often are...the practice of automatically collecting customers’ account numbers, and then making those customers repeat the numbers to an agent when they finally connect to one -- my top IVR gripe"
"Voice calls must be able to recruit data channels, and vice versa. That way, an agent could attach an IM session to your voice call and push you the URL in real-time chat. It might even be appropriate to extend the data session with screen sharing, so the agent can watch and assist. If things still don’t work out and the whole matter must be referred to someone else, you’d like to be able to initiate voice or data communication -- or both -- in a context-preserving way."
Via, Rethinking customer service. This points to XBL2 and an effort to provide "...a declarative format for applications and user interfaces...based on an existing application/UI format, such as Mozilla's XUL, Microsoft's XAML, Macromedia's MXML or Laszlo Systems' LZX..."
Beyond interactive voice response "English’s hope, he tells Inskeep in the interview, is that companies will admit how infuriating their systems often are...the practice of automatically collecting customers’ account numbers, and then making those customers repeat the numbers to an agent when they finally connect to one -- my top IVR gripe"
"Voice calls must be able to recruit data channels, and vice versa. That way, an agent could attach an IM session to your voice call and push you the URL in real-time chat. It might even be appropriate to extend the data session with screen sharing, so the agent can watch and assist. If things still don’t work out and the whole matter must be referred to someone else, you’d like to be able to initiate voice or data communication -- or both -- in a context-preserving way."
Via, Rethinking customer service. This points to XBL2 and an effort to provide "...a declarative format for applications and user interfaces...based on an existing application/UI format, such as Mozilla's XUL, Microsoft's XAML, Macromedia's MXML or Laszlo Systems' LZX..."
Model Driven, Semantic Web Enabled, Science Commons
Semantic Web eyed for life sciences data "The Semantic Web involves a concept in which data from multiple sources and ontologies can be integrated into a single information space. Experiment design automation (XDA) software vendor Teranode, which focuses on software for life sciences, plans to collaborate with Science Commons to build a neurology repository for the Semantic Web."
NeuroCommons is part of the ScienceCommons project, it is going to provide a database and annotations of scientific data in (presumably) RDF.
Teranode explains with their XDA product, why model driven and why the semantic web.
A related post, via Etymon.com Federated Databases in Science "The astronomy, chemistry, and geospatial communities were active well over a decade ago in collaborating with information scientists on federated databases through various open standards. Molecular biology is a field that currently has considerable needs in this area, stimulated by the Human Genome Project. Developing common standards through consensus is of course not a technological solution. The Web is successful because it exploits the relationships among a huge number of people making individual judgements that only people can make. Even the Semantic Web, if it ever has a chance of working, would have to depend on a very large base of common metadata standards, and that can only result from the slow process of people coming together and agreeing. There are many things that information technology cannot do on its own. The semantic integration of knowledge still remains a human activity."
NeuroCommons is part of the ScienceCommons project, it is going to provide a database and annotations of scientific data in (presumably) RDF.
Teranode explains with their XDA product, why model driven and why the semantic web.
A related post, via Etymon.com Federated Databases in Science "The astronomy, chemistry, and geospatial communities were active well over a decade ago in collaborating with information scientists on federated databases through various open standards. Molecular biology is a field that currently has considerable needs in this area, stimulated by the Human Genome Project. Developing common standards through consensus is of course not a technological solution. The Web is successful because it exploits the relationships among a huge number of people making individual judgements that only people can make. Even the Semantic Web, if it ever has a chance of working, would have to depend on a very large base of common metadata standards, and that can only result from the slow process of people coming together and agreeing. There are many things that information technology cannot do on its own. The semantic integration of knowledge still remains a human activity."
Friday, December 09, 2005
Graphical Batch Files
AutoMate This is a pretty interesting application that basically provides similar functionality that OSX Automator provides. While you can create customized tasks in VBA it has lots of interesting inbuilt functionality manipulating Excel, FTP, terminal emulation, keyboard and windows manipulation.
Related, but much simpler is a piece of software called, AutoIt which was "...initially designed for PC "roll out" situations to reliably configure thousands of PCs, but with the arrival of v3 it has become a powerful language able to cope with most scripting needs." Basically allowing simple automation of window, mouse and keyboard events.
Related, but much simpler is a piece of software called, AutoIt which was "...initially designed for PC "roll out" situations to reliably configure thousands of PCs, but with the arrival of v3 it has become a powerful language able to cope with most scripting needs." Basically allowing simple automation of window, mouse and keyboard events.
Thursday, December 08, 2005
One Way Java is Better than Ruby
The unbridled humanity of APIs "But I think the Java guy has a point: 78 methods on your list objects isn't good. Less methods is good. Unless the result is stupid. Now, let's be honest here, Java is stupid. Dumb, idiotic, maybe written by people who aren't programmers; I just don't know how else to make sense of it. list.get(list.size() - 1) should be embarrassing. list.last or list[-1]? I think [-1] reads well enough, and fits into a very elegant set of functionality involving slices and whatnot. But I also think list.last is entirely justifiable. OTOH, list.get(0) isn't embarrassing, so list.first isn't as compelling."
"Maybe an interesting parallel is 0 vs. 1 indexing. 1 clearly seems more humane. I personally count starting from 1. I'm naturally inclined to index from 1. Languages go both ways on the choice...Of course Smalltalk indexes from 1, so no one gets everything right."
Humane Interfaces "Part of the reason this argument could go on forever is that Ruby’s Array is both an example of arguments for Humane design, and arguments against it...java.util.List isn’t really a shining example of good interface design either...Having two otherwise equivalent ways to perform the same operation is bad user-interface design, and it’s bad library interface design, because the existence of the synonyms actually adds to your cognitive load by making you choose between them."
Also, Why Ruby Shouldn’t Be Your Next Programming Language (Maybe).
"Maybe an interesting parallel is 0 vs. 1 indexing. 1 clearly seems more humane. I personally count starting from 1. I'm naturally inclined to index from 1. Languages go both ways on the choice...Of course Smalltalk indexes from 1, so no one gets everything right."
Humane Interfaces "Part of the reason this argument could go on forever is that Ruby’s Array is both an example of arguments for Humane design, and arguments against it...java.util.List isn’t really a shining example of good interface design either...Having two otherwise equivalent ways to perform the same operation is bad user-interface design, and it’s bad library interface design, because the existence of the synonyms actually adds to your cognitive load by making you choose between them."
Also, Why Ruby Shouldn’t Be Your Next Programming Language (Maybe).
Time you enjoy wasting, was not wasted
- "TestDriven.NET makes it easy to run unit tests with a single click, anywhere in your Visual Studio solutions. It supports all versions of Microsoft Visual Studio .NET meaning you don't have to worry about compatibility issues and fully integrates with all major unit testing frameworks including NUnit, MbUnit, & MS Team System." Discussion on using TestDriven.NET with Express. Via James.
- Obie Fernandez from ThoughtWorks on Ruby and the Semantic Web "Java Doesn’t Work for Ontologies...Polymorphism in RDF and OWL is very different than in Java which makes for admittedly clumsy API...Why Deep Integration Enhances Programming with Ruby...Gives Ruby programmer powerful inferencing capabilities against the underlying RDF database". Related to Scripting the Semantic Web and Northrop Buys Tucana and Continues Kowari.
- OO in One Sentence: Keep It DRY, Shy, and Tell The Other Guy Available here. Covers how OO methods are not function calls but message passing.
- Beyond Story Cards: Agile Requirements Collaboration "Using this idea of "latest responsible moment," we postpone creating our detailed requirements specification. In fact, we can do the specification for a story at the same time as we actually implement the story. (I'll be talking about that more in a little bit.) We don't need to know every detail about a feature ahead of time. Going into the level of detail is expensive, and if (or when) our plans change, that's wasted effort."
Wednesday, December 07, 2005
Making AJAX cool
Why Ajax Sucks (Most of the Time) "For new or inexperienced Web designers, I stand by my original recommendation. Ajax: Just Say No."
"Ajax breaks the unified model of the Web and introduce a new way of looking at data that has not been well integrated into the other aspects of the Web. With ajax, the user's view of information on the screen is now determined by a sequence of navigation actions rather than a single navigation action.
Navigation does not work with ajax since the unit of navigation is different from the unit of view. If users create a bookmark in their browser they may not get the same view back when they follow the bookmark at a later date since the bookmark doesn't include a representation of the state of the content on the page.
Even worse, URLs stop working: the addressing information shown at the top of the browser no longer constitutes a complete specification of the information shown in the window."
"Ajax breaks the unified model of the Web and introduce a new way of looking at data that has not been well integrated into the other aspects of the Web. With ajax, the user's view of information on the screen is now determined by a sequence of navigation actions rather than a single navigation action.
Navigation does not work with ajax since the unit of navigation is different from the unit of view. If users create a bookmark in their browser they may not get the same view back when they follow the bookmark at a later date since the bookmark doesn't include a representation of the state of the content on the page.
Even worse, URLs stop working: the addressing information shown at the top of the browser no longer constitutes a complete specification of the information shown in the window."
Tuesday, December 06, 2005
A little light I/O
Comparing Two High-Performance I/O Design Patterns "It is clear from the charts that C++ is still the preferable approach for high performance communication solutions, but Java on Linux comes quite close. However, the overall Java performance was weakened by poor results on Windows. One reason for that may be that the Java 1.4 nio package is based on select()-style API. Ð It is true, Java NIO package is kind of Reactor pattern based on select()-style API (see [7, 8]). Java NIO allows to write your own select()-style provider (equivalent of TProactor waiting strategies). Looking at Java NIO implementation for Windows (to do this enough to examine import symbols in jdk1.5.0\jre\bin\nio.dll), we can make a conclusion that Java NIO 1.4.2 and 1.5.0 for Windows is based on WSAEventSelect () API. That is better than select(), but slower than IOCompletionPortÕs for significant number of connections. . Should the 1.5 version of Java's nio be based on IOCompletionPorts, then that should improve performance. If Java NIO would use IOCompletionPorts, than conversion of Proactor pattern to Reactor pattern should be made inside nio.dll. Although such conversion is more complicated than Reactor- >Proactor conversion, but it can be implemented in frames of Java NIO interfaces. (this the topic of next arcticle, but we can provide algorithm). At this time, no TProactor performance tests were done on JDK 1.5."
Available in Java and C++ at Terabit.
Available in Java and C++ at Terabit.
Sunday, December 04, 2005
Links to Share and Enjoy
Another list of links:
- XML 2005: Tipping Sacred Cows "Which brings us to one of our sacred cows: for decades we've had SQL for relational databases, and soon we'll have XQuery for general XML, and SPARQL for RDF...what if it was possible to construct a generalized query language, loosely coupled enough to work with any underlying data model? The mathematical basis for this was monoids. The presentation didn't actually define this fairly abstract term, only skipping from trivial examples like
or to a fully worked representation of a generalized query. Erik's dynamic presentation style is such that I was not able to copy down the full example before he had moved on to the next slide. Whatever the details, it's a valuable topic in that it gets listeners to question their assumptions and see in new ways." - Dabble combines the best of group spreadsheets, custom databases, and intranet web applications into a new way to manage and share your information online. A lot of the same conversations in Agile databases seem to occuring - has related to functionality, data integrity, migration (string to first name, last name, for example), data types and the like. Includes a blog and demo (movie). Merging is coming in version 2.0 and it's RAM based. It just cries out for an RDF data model. Via, Dabble is Bloody Brilliant.
- Problems with the $100 laptop "The time will certainly come when the appropriate tool to promote economic development will be a laptop produced very inexpensively in large volume. Before that point it will be necessary to implement systems that provide infrastructure which the laptop will need, in addition to producing tangible economic benefits for their users. OLPC is to be commended for raising issues and focusing attention, and for posing some technological challenges in a highly visible way...large sums of money are to be committed to the project in advance to fund manufacturing in deals where the customers are government ministries and not the end users." Also, $100 laptop.
- Two interesting articles: Breaking The Quality–Speed Compromise "The most important thing we can do to break the compromises we impose on customers is to move testing forward and put it in-line with (or prior to) coding. Build suites of automated unit and acceptance tests, integrate code frequently, run the tests as often as possible. In other words, find and fix the defects before they even count as defects." and Is Agile Software Development Sustainable? "So if agile practices are a “disruptive technology” compared to traditional software development processes, then it would be quite in character for them to start by addressing small systems."
- Exploratory Testing on Agile Projects Can Be a Good Fit "Why should agile teams do exploratory testing?: "Because an agile development project can accept new and unanticipated functionality so fast, it is impossible to reason out the consequences of every decision ahead of time. In other words, agile programs are more subject to unintended consequences of choices simply because choices happen so much faster. This is where exploratory testing saves the day. Because the program always runs, it is always ready to be explored.""
- Matthew De George on Cranky Middle Manager. Explaining how to apply market economies to management and more. Forgive Matthew for his hierachical view, it's all about graphs of course. I'm a bit slow in finding this.
- A different way to vommit in another yearly milestone: Tiger Moth Joy Flights.
Saturday, December 03, 2005
Best things in development are free
One of the key differences between Java and .NET development is cost. To get the right Microsoft solution costs thousands of dollars. And what you get is something very different to an IntelliJ, NetBeans or Eclipse. You get vendor integration (or lock in if you prefer) and competition against the community. It may appear attractive to some but it seems odd to me to actively fight integration into existing, open solutions. It does seem though that the open side is winning (I wish Sun had've done the same thing when choosing a logging API).
Unit testing and source control are obvious ones. It's amazing to see that their entry level IDE does not come with this. There are great free solutions in NUnit and MbUnit. MbUnit is especially cool (released recently) allowing all sorts of built in test fixtures. And there's always Ankh for Subversion integration. There are lots of alternatives.
A summary from a Microsoft developer: Hey, Shareholders! VS 2005 is *Fantastic* and our Developers Love Microsoft! "I might wander in early on Monday to meander through the crowds celebrating the big Visual Studio launch. But my heart is heavy that we shoveled what we could together and Won't Fix-ed this release out the door. Microsoft has just opened a very big door to competition in the IDE space. Or at least towards people jealously holding onto VS 2003 and saying, "CLR 2.0? Screw that! The last time I tried to use generics my machine locked up!" Big freakin' mistake. Microsoft should be ashamed."
Unit testing and source control are obvious ones. It's amazing to see that their entry level IDE does not come with this. There are great free solutions in NUnit and MbUnit. MbUnit is especially cool (released recently) allowing all sorts of built in test fixtures. And there's always Ankh for Subversion integration. There are lots of alternatives.
A summary from a Microsoft developer: Hey, Shareholders! VS 2005 is *Fantastic* and our Developers Love Microsoft! "I might wander in early on Monday to meander through the crowds celebrating the big Visual Studio launch. But my heart is heavy that we shoveled what we could together and Won't Fix-ed this release out the door. Microsoft has just opened a very big door to competition in the IDE space. Or at least towards people jealously holding onto VS 2003 and saying, "CLR 2.0? Screw that! The last time I tried to use generics my machine locked up!" Big freakin' mistake. Microsoft should be ashamed."
Subscribe to:
Posts (Atom)