LinuxWorld
LinuxWorld Conference and Expo August 4-7, 2008 Call for papers open until Feb. 22

OpenPipeline seeks to ease document prep for search

Enterprise search vendor Dieselpoint is behind a new open-source project centering on a document "pipeline" -- or as the Chicago company's CEO, Chris Cleveland, puts it, "all the boring stuff you need to make enterprise search work."

Related links

No results were found for your search.

Your query is too restrictive.
You might want to try: data center

Enterprise search implementations often cover an array of document sources and components; pipelines allow companies to standardize the processing of information before it gets pushed into a search-engine indexer.

"We're connecting the crawler companies to the text analytic companies to the search engine companies," Cleveland said.

Dieselpoint was having trouble integrating its own pipeline with third-party document analyzers and content connectors, and has open-sourced it as a basis for the project, which is dubbed OpenPipeline.

Its Web site is scheduled to open to the public on Monday, and a fully functional version of the software will be downloadable under the Apache 2.0 license. It is available under a commercial license as well, according to the site.

The software features a point-and-click user interface and provides a number of connectors, including Web and SQL crawlers. It also supports a number of commercial connectors for products such as SharePoint, Exchange and a number of portals.

Dieselpoint is pursuing the project both to make bigger, more complex implementations easier and in hopes that it will draw some customers to its search engine.

"The single biggest barrier to adoption of enterprise search is doing integration," Cleveland said. "Of course, it means enormous consulting engagements, so it's a source of revenue for the industry, but it's a deterrent."

While major search vendors have pipelines, they are "all proprietary and all closed," he said.

A number of other vendors and consultants have signed on to the effort's advisory board. They include Alias-i, Applied Relevance and Raritan Technologies. Cleveland is anticipating more companies will join soon.

Conceptually, an open-source pipeline makes sense for the industry on the whole "because each component is worthless on its own," he suggested.

1 | 2 |  Next >

React: Give us your thoughts on the issues here.
Use this form to start a public discussion with other Linux World users on this article.
Log In | Register for an account (Why you should)

Note: Register to have your user name appear; otherwise your comment will show up as "Anonymous."

*Anonymous comments will only appear once they are approved by the moderator.

Newsletter sign-up

Sign up for one of Network World's newsletters compliments of Linux World

Linux & Open Source News Alert
Web Applications Alert
Video & Podcast Alert
Security: Threat  Alert
Virtualization Alert

Email Address: