Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Thursday, May 8, 2014

"Predictive" Coding and the Naked Emperor

I've always been suspicious of the claims of the "predictive coding" zealots. Every year it seems that there is a new buzzword in the field of  Ediscovery. Technology Assisted Review - Check! Information Governance - Check! Cloud Computing - Check. Big Data - Check.  

My friend John Martin did a magnificent job of distilling some of my concerns in his blog entry that can be found right here.

The Emperor has No Clothes - and PC Can't See Image-Only Documents

There are several parallels between predictive coding (AKA technology assisted review) and Hans Christian Andersons' tale, "The Emperor's New Clothes." In the story, two weavers tell the emperor they will make him a suit of clothes that will be invisible to those people who are unfit for their position, stupid, or incompetent. None of the emperor's subjects want to admit to those deficiencies so the emperor parades around with no clothes on until a child states the obvious - the emperor has no clothes.
Predictive Coding - do not see any evidence hereIn the case of predictive coding, its advocates have touted the efficacy of their approaches in white papers, blogs, and social media postings, and have practically created a separate industry to host conferences promoting the wonders of predictive coding. Few people want to ruin the moment or buck the trend by pointing out what is obvious when one considers the technology underlying predictive coding - it is completely dependent on having text to analyze.  It will absolutely fail to analyze documents for which there is no text, and will do a miserable job where the text is of poor quality.
This might be just an esoteric debating point if virtually all documents had associated text. However, in some industries like oil & gas, half or more of some collections will be engineering drawings and schematics that were output to image-only PDF for distribution and use by those who don't have the software licenses needed to view the documents in their original file formats.
Predictive Coding 100 percent right 20 percent of the timeIn practically all industries it is common practice to develop documents in one application like Word and then, once finalized, distribute them as image-only PDF so they can be viewed on a variety of devices and so recipients can't easily change the content. In one collection we analyzed, only 20% of the PDFs had associated text. Even if predictive coding were 100% effective, the most it could classify would be 20% because it literally cannot "see" the 80% without text. If in fact predictive coding has a recall rate of 70-80% of what it can see, that would mean that predictive coding would have identified 14 to 16% of the total PDFs (70% x 20% = 14% or 80% of 20% = 16%). By contrast, BR's visual classification technology classified 100% of them.
PDFs will potentially be among the most relevant file types in a collection because that is the format used to distribute information within and among groups of people within an organization, and among organizations. Note that even if in some unique e-discovery settings predictive coding is acceptable, the text-restriction failing of predictive coding will be fatal for broader information governance purposes.
So... if you're going to use predictive coding, at the very least measure what PC doesn't "see." If you're planning on using PC for information governance purposes, make sure that the organization doesn't mind not classifying a potentially significant percentage of its documents.

Thursday, October 11, 2012

ISV Takes Fresh Look at Recognition

DOCUMENT IMAGING REPORT
September 28th, 2012



There is no question that ISVs are currently trying to go where no capture software has gone before—in terms of applying automatic recognition technology to documents. Over the past couple issues, we’ve covered topics like artificial intelligence, semantics, and advanced analytics, all designed to take applications beyond the capabilities of current capture software. However, an appropriately named start-up out of Germantown, TN (near Memphis), may have beaten many established capture players to the mark.
Using a combination of advanced pattern recognition and computer vision, BeyondRecognition recently completed a project in which it successfully indexed 2.3 billion images, which originally contained no meta data. Yes, working as a contractor for an energy company, BeyondRecognition was presented with 27,000 CDs and DVDs full of images that covered a timeframe of roughly 90 years and were created in different locations across the world.
 "There were no boundaries separating the scanned images,” said John Martin, founder and CEO of BeyondRecognition. “The only thing we knew was that the documents on the discs in the front of each box were scanned before the documents on the discs in the back. Our client wanted to be able to mine the data on all these documents.”
According to Martin, he looked at everything available on the market for accomplishing the task at hand. “I looked at traditional OCR applications, but they were a bust—even with voting engines, which would just have made the process five or six times longer,” he said. “Even if we could have applied full-text OCR, conventional search engines could not do the things we wanted. In addition, traditional relational databases couldn’t support the millions of many-to-many relationships we had to set up.”
To build the application that eventually became of cornerstone of BeyondRecognition, Martin applied a process he called “negative learning.” “Basically, I started with no pre-conceived ideas about how automatic recognition was being done,” he said. “Instead, I tried to take the position that if I were to build a recognition solution with tools available today, how would I approach it?”
The guts of the solution
Martin started with the basic premise that computers are good at working with numbers. “Based on that, we were able to identify one of the key flaws of traditional OCR—it attempts to read characters like humans do,” he said. “It goes left to right, top to bottom, first page to last.
“And it attempts to recognize each character individually. Think about that from a statistical standpoint. On each page, a single lowercase character might appear 40 to 50 times. So, on a million pages, that character could show up as many as 50 million times. Basically, with traditional OCR, you’re giving software 50 million chances to get it wrong. Statistically, that means, it’s certainly going to make at least one mistake.”
BeyondRecognition puts each image through a process it calls “scraping.” “We literally rip the images apart into glyphs,” said Martin. “Those glyphs include not only the characters on a page, but also things like logos, staple holes, check boxes, signatures, and even specks of dirt. On average, we produce about 1,500 glyphs per page.
“We then run the glyphs through a normalization process before grouping them. The normalization involves accounting for orientation by rotating each glyph 720 times in half-degree increments. This way, direction doesn’t matter, and neither does size. After it’s normalized, if a glyph is 99% similar or greater to other glyphs, they are placed in the same cluster.”
What happens next is a bit confusing, but it basically involves identifying these clusters as sets of characters. This is accomplished at least partially by identifying the glyph in a cluster that most exactly resembles a known character and then plugging that character into a word that is checked against a global dictionary. There are also statistical formulas incorporated regarding how often a particular character should show up in a set of documents.
“We average the results of all that, and if it comes back above a 99% confidence rating that the glyph represents a specific character, we presume it to be true,” said Martin. “Then, because our software tracks the location of each glyph it creates, we can identify all the glyphs in that particular cluster as being that particular character.”
Of course, not every glyph comes back at a 99% confidence level. To account for this, Martin showed us a process called “Word QC.” In the example he showed “lockbox” was not recognized as a valid word in the global dictionary, so it was highlighted on the screen. The statistics said it was one of several million suspect words (in a large set of documents). Merely confirming that “lockbox” was a valid word had a cascading effect that helped validate other glyphs as characters. The result was that with a single keystroke, 900,000 suspect words were eliminated. “That’s the type of stuff, offshore keyers are being paid to correct on a word-by-word basis,” said Martin.
BeyondRecognition has the ability to output searchable PDF files, as well as what it calls an “XPDF” file. “Basically, this is a cross-reference file, which includes a coordinate point for every word and numeric sequence pulled off a page,” said Martin. “It’s a great tool for redaction applications, for example.
“We have a customer using it to redact expressions like Social Security and phone numbers. They can achieve this at a rate of 600,000 redactions per hour.”
Because BeyondRecognition works with glyphs, it is able to handle multiple languages—even mixed within a single document set.
It also has the ability to do document clustering based on the layout and content of images. “This is a great tool for automatically routing files to the right process,” said Martin. “We have a BPO that utilizes that element of our technology to help it allocate its resources more effectively.
“In the discovery space, document clustering can help users eliminate a lot of documents they’re not going to need. For example, in a labor case, it enables them to quickly identify tax and marketing documents that won’t be relevant.”
This type of document classification also has obvious potential in markets like mortgage and patient records processing, where several types of documents are often mixed in a single file.
Rules can be set up within BeyondRecognition’s application for extracting specific data fields and tables. Extraction can be applied to structured and semi-structured documents.
Rules can also be set up around the glyphs to eliminate background noise such as watermarks. Martin showed an example of image enhancement being applied to documents created through carbon-paper duplication.
Martin said the speed of BeyondRecognition’s software depends on the number of CPUs being utilized. “A 30-core server can process up to a million pages per day,” he said. “An 80-core can do 5 million, and a 160-core, about 10 million.
To date, BeyondRecognition has offered its technology solely as a service. “We’ve done several dozen projects, including several different types,” said Martin. “We’re currently working on developing an appliance that can be run behind a customer’s firewall.”
Martin said that BeyondRecognition’s technology belongs entirely to his company. “There are some patents around it, as well as some we’re applying for,” he said. “We have 15-16 man years worth of development in this.”
Martin concluded by echoing the theme we thought was prevalent at the recent Harvey Spencer Associates Capture Conference, regarding next generation document capture. “This is a solution for big data,” he said. “If you don’t know what you have, you certainly can’t decide what’s relevant.”
For more information: http://www.beyondrecognition.net

Sunday, September 2, 2012

Willie Wonka, Big Data, Pure Imagination


Hold your breath, Make a wish, Count to Three

Come with me …And you'll be
In a world of Pure imagination
Take a look, And you'll see Into your imagination

We'll begin
With a spin…Traveling in
The world of my creation…What we'll see
Will defy….Explanation





Maybe you thought that BeyondRecognition was merely the greates automatic “Visual Document Clustering” New Tool for Big Data. Well you'd be correct yet wrong. BeyondRecognition's powerful graphical engine can provide never before dreamed of  levels of image enhancement. 


BeyondRecognition's Visual-Similarity Clustering automatically processes and groups documents together for document boundary detection and document type classification, regardless of source and format — seamlessly processing native electronic files and scanned documents
Automatic Visual Document Clustering means that visually similar pages are gathered based on their graphical, rather than textual, content.  This avoids the errors normally encountered in extracted or generated text and leverages non-text graphical elements such as logos, form elements and other objects to greatly improve accuracy.

About BeyondRecognition

BeyondRecognition is a "textnology" company that has developed unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. Disclosure of further details is being deferred until one or more patents on the process are filed. BeyondRecognition is working with a select number of companies in the electronic discovery and document management industries. 
For more information, visit www.BeyondRecognition.net.

About Focus Data Management
FDM is the sales and marketing arm of BeyondRecognition. With offices on both coasts FDM is available to help you customize your Big Document Solutions using the power of BeyondRecognition. Contact us at 804.690.0010 or 562.822.7141

Tuesday, July 31, 2012

Dr. Stephen V. Rice Joins BeyondRecognition Advisory Board





Germantown, TN (July31, 21012). John Martin, founder and CEO of BeyondRecognition, LLC, announced today that Stephen V. Rice, Ph.D., has agreed to join BeyondRecognition’s Advisory Board and to consult with BeyondRecognition.  Martin indicated that, “One of the major functionalities provided by BeyondRecognition is our ability to extract textual and other non-textual glyphs from document images.  Dr. Rice’s scientific background uniquely qualifies him to help us to create objective measures of the accuracy of that conversion process at the character, word, and significant-word level.”

For five years, Dr. Rice conducted the first large-scale independent evaluations of commercial optical character recognition (OCR) systems while at the Information Science Research Institute of the University of Nevada, Las Vegas (UNLV). As part of that work he developed sequence comparison algorithms to measure OCR accuracy. He is author of the classic book, “Optical Character Recognition: An Illustrated Guide to the Frontier,” which is essential reading for anyone involved in developing OCR or CAPTCHA systems.

Rice noted that, “BeyondRecognition has taken a fresh approach to the challenge of extracting text from document images. They have developed several innovative technologies for document conversion and retrieval.  I look forward to assisting them as they continue to break new ground in this area.”

Martin continued, “For the past year we have been building out our code base and our infrastructure and will be looking to Dr. Rice to help us develop statistically-sound performance measures to evaluate our performance. For example, next month we plan to benchmark our system on the approximately 30 million page images of tobacco litigation documents obtained from the Text Retrieval Conference (TREC) sponsored by the National Institute of Standards and Technology (NIST). Our goal is to perform the text conversion on the 30 million pages, globally edit the text, index it and be able to perform sub-second retrieval on any page in the collection within 72 hours, start to finish. We want valid, reliable metrics to use to report the results.”

As to the significance of the document collection, Martin observed, “The TREC Tobacco documents have been used by the TREC Legal Track to gauge the efficacy of various text retrieval systems and methodologies, despite known issues of using inaccurate OCR. The studies compared the results from queries or processing to the documents identified by manual reviews, and the results have been used to argue in favor of ‘predictive coding’ or ‘technology-assisted review.’ We believe that the new, more accurate text output by BeyondRecognition will let researchers show how the early studies may have understated the effectiveness of technology-assisted review.”

About Dr. Rice

Dr. Stephen V. Rice consults as a computer scientist and software engineer with expertise in algorithms, computer audio, computer simulation, database systems, pattern recognition, programming languages, and related areas. He was a computer science professor at the University of Mississippi, was chief software engineer at the UNLV Information Science Research Institute, and is Founder and CTO of Comparisonics Corporation. He served on the board of advisors of the Federal Intelligent Document Understanding Laboratory of the U.S. Central Intelligence Agency. 

For more about Dr. Rice, see http://www.stephenvrice.com

About BeyondRecognition

BeyondRecognition has developed unique character, word, and document attribute recognition and extraction capabilities for analyzing image-based documents. Its glyph clustering and cataloging approach enables rapid, globally-editable text recognition with accuracy rates far beyond traditional OCR. BeyondRecogntion also clusters documents based on visual similarity and permits location-based, cluster-specific data element extraction for coding or abstracting data elements from the documents. Clustering by document type permits prioritized data element extraction using the powerful graphical user interface to highlight zones, and to write and instantly test and verify extraction rules.

Although nominally a “startup,” the principal technologists at BeyondRecognition have been working in the fields of document conversion, electronic evidence forensics and processing for decades. CEO John Martin was previously a founder of Cricket Technologies and RedFile LLC.
For more information, visit www.BeyondRecognition.net

Thursday, July 26, 2012

Obama investing 200 million in big data research projects

Thanks to new approaches for processing, storing and analyzing massive volumes of multi-structured data – such as Hadoop and MPP analytic databases — enterprises of all types are uncovering new and valuable insights from Big Data everyday.

graph courtesy of wikkibon.org

Wednesday, June 13, 2012

MAY BRINGS RICH C-LEVEL EXPERIENCE IN INDUSTRY AND PHILANTHROPY TO HIGH-TECH STARTUP BEYONDRECOGNITION


Germantown, TN: (May 30, 2012). John Martin, founder and CEO of BeyondRecognition, LLC, today stated that, ”BeyondRecognition is pleased to announce that Ken May will be providing business development guidance for BeyondRecognition as it pushes its innovative image-based document analysis technology into key markets like mortgage and loan processing, and the oil and gas industries.” BeyondRecognition’s breakthrough integrated workflow enables companies to obtain actionable intelligence from image-based and electronic format documents at a fraction of the cost associated with manually reviewing and abstracting paper files and often with higher accuracy and reliability.

Martin continued, “BeyondRecognition’s core competencies lie in document processing and analysis, and Ken brings an incredible wealth of experience managing FedEx Kinkos, one of the largest and most wide-spread document copying and handling operations in the world, as well as planning and managing some of the most highly automated decision-support systems in the world. He also has a wealth of C-level contacts at companies across America from his many years of service as Chairman of the National Board of Trustees for the March of Dimes. We look forward to being able to capitalize on his rich experience, energy, and industry knowledge.”

Ken May commented, “I have had the opportunity over the years to review many exciting technologies at all sorts of start-ups and emerging market leaders, but I was especially struck at how innovative BeyondRecognition’s technology is and at the incredible value it offers companies that are faced with needing to analyze and process large volumes of paper-based records. This is particularly true in industries like home loan processing where the documents in the underlying files are typically not all or even mostly electronic. The need to process existing back files of loan documents and to eventually automate the new loan initiation process represents an enormous potential. I look forward to helping spread the message about this important new technology.”

About Ken May

Beginning as a manager of hub operations for FedEx in 1982, May served in various management positions, becoming VP, Global Operations Scheduling and Control in 1996. He then served as Sr. VP, Air-Ground and Freight Services 1997 to 1999, was Sr. VP US Operations from 1999 to 2004, COO at FedEx-Kinko’s Office and Print Centers from 2004 to 2006, and was President and CEO at FedEx-Kinko’s Office and Print Centers from 2006 to 2008. 
May served as Chairman of the National Board of Trustees at the March of Dimes from 2007 to 2011, and was President of ES3, LLC, the third-party logistics subsidiary of C&S Wholesale Grocers, the eighth-largest privately held company in the US by revenue from 2010 to 2011. From 2011 to 2012 he was President and COO at Krispy Kreme Doughnuts.

May has been a director of PF Chang’s China Bistro since May 2007, and serves on the Board of Directors of Greystone Medical Group. 

For more about Ken May, see http://en.wikipedia.org/wiki/Ken_may.

About BeyondRecognition

BeyondRecognition has developed unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. Its glyph clustering and cataloging approach enables rapid, globally-editable text recognition with accuracy rates far beyond traditional OCR. BeyondRecognition also clusters documents based on visual similarity and permits location-based, cluster-specific data element extraction for coding or abstracting data elements from the documents. Clustering by document type permits prioritized data element extraction using the powerful graphical user interface to highlight zones, and to write and instantly test and verify extraction rules.

Although nominally a “startup,” the principal technologists at BeyondRecognition have been working in the fields of document conversion, electronic evidence forensics and processing for decades. CEO John Martin was a founder of Cricket Technologies, LLC and RedFile LLC.

For more information, visit www.BeyondRecognition.net

Wednesday, June 6, 2012

Unlocking Paper Based Intelligence with Disruptive Technology

"You want your documents to be searchable - not laughably searchable . . . " John Martin



When confronted with massive amounts of unstructured data and the need to access the business intelligence locked inside that data the options before today were expensive, required massive human intervention, were extremely time consuming and most troubling, very ineffective. John Martin loves disruptive technology. I love how John Martin thinks.

John's latest game changing software BeyondRecognition is being referred to as a "Big Data innovator" by several of the " Big Four" accounting  firms. The tool was originally built  for a  company that  needed to extract key information from a 30+ year old scanned paper document set of 2.3 BILLION pages for a due diligence effort.  In the energy sector, Beyond Recognition's  glyph clustering technology makes it possible to search for symbols 
used on maps to indicate things like radioactive wells, salt-water wells or API number codes.

As a result of this disruptive new technology Martin notes , "We're seeing a great deal of interest in this approach in the mortgage and energy sectors. The mortgage industry in particular typically has a relatively finite number of documents in loan files supporting the loan decisions, with definable types of data being of interest on each type of document. Our process could greatly lower the cost of tracking all those data elements during loan initiation, or to quality control the file for audit or sale purposes."




Essentially BeyondRecognition's  unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. In Plain english  BR allows clients to extract very valuable business intelligence from scanned and digitized files fast, accurately and in a cost-efficient manner. In a press release Barbara Johnson, CFA, former executive of USAA Federal Savings Bank, serving as Chief Credit Officer and Senior Vice President of Real Estate Lending Services and now a Principal with Saccadent, a Financial Consulting firm, has reviewed the clustering and data extraction capabilities and offered the comment that, "In today's environment the ability to extract, utilize and match data across a variety of documents is incredibly powerful. This type of technology offers the promise of significantly decreasing the time and cost to process a thoroughly compliant loan from application through origination, audit, sale and servicing. An automated system to confirm all the critical items match throughout the process and are in the appropriate format and location on all documents would be invaluable. Lending is a document-rich industry and the time is perfect for this type of technology."

Another intriguing aspect is that  BR solution is language agnostic automatically recognizing 40+ languages interspersed throughout any data set with no up-front programming required and performs at 99.5% word accuracy on first pass unassisted review. The solution can scale to meet customer requirements between 500k - 50m pages per day regardless of the legibility of the images. In fact using BR Adaptive Image Enhancement techniques restorion of  poor quality document images to much improved legibility is a seemless byproduct 


For additional information or to request a demo  contact John Martin at John AT beyond recognition DOT net
or Michael Mulcahy at Michael@focusdata-mgt.com or by phone at 562 546-2465