Showing posts with label document technology. Show all posts
Showing posts with label document technology. Show all posts

Thursday, May 8, 2014

"Predictive" Coding and the Naked Emperor

I've always been suspicious of the claims of the "predictive coding" zealots. Every year it seems that there is a new buzzword in the field of  Ediscovery. Technology Assisted Review - Check! Information Governance - Check! Cloud Computing - Check. Big Data - Check.  

My friend John Martin did a magnificent job of distilling some of my concerns in his blog entry that can be found right here.

The Emperor has No Clothes - and PC Can't See Image-Only Documents

There are several parallels between predictive coding (AKA technology assisted review) and Hans Christian Andersons' tale, "The Emperor's New Clothes." In the story, two weavers tell the emperor they will make him a suit of clothes that will be invisible to those people who are unfit for their position, stupid, or incompetent. None of the emperor's subjects want to admit to those deficiencies so the emperor parades around with no clothes on until a child states the obvious - the emperor has no clothes.
Predictive Coding - do not see any evidence hereIn the case of predictive coding, its advocates have touted the efficacy of their approaches in white papers, blogs, and social media postings, and have practically created a separate industry to host conferences promoting the wonders of predictive coding. Few people want to ruin the moment or buck the trend by pointing out what is obvious when one considers the technology underlying predictive coding - it is completely dependent on having text to analyze.  It will absolutely fail to analyze documents for which there is no text, and will do a miserable job where the text is of poor quality.
This might be just an esoteric debating point if virtually all documents had associated text. However, in some industries like oil & gas, half or more of some collections will be engineering drawings and schematics that were output to image-only PDF for distribution and use by those who don't have the software licenses needed to view the documents in their original file formats.
Predictive Coding 100 percent right 20 percent of the timeIn practically all industries it is common practice to develop documents in one application like Word and then, once finalized, distribute them as image-only PDF so they can be viewed on a variety of devices and so recipients can't easily change the content. In one collection we analyzed, only 20% of the PDFs had associated text. Even if predictive coding were 100% effective, the most it could classify would be 20% because it literally cannot "see" the 80% without text. If in fact predictive coding has a recall rate of 70-80% of what it can see, that would mean that predictive coding would have identified 14 to 16% of the total PDFs (70% x 20% = 14% or 80% of 20% = 16%). By contrast, BR's visual classification technology classified 100% of them.
PDFs will potentially be among the most relevant file types in a collection because that is the format used to distribute information within and among groups of people within an organization, and among organizations. Note that even if in some unique e-discovery settings predictive coding is acceptable, the text-restriction failing of predictive coding will be fatal for broader information governance purposes.
So... if you're going to use predictive coding, at the very least measure what PC doesn't "see." If you're planning on using PC for information governance purposes, make sure that the organization doesn't mind not classifying a potentially significant percentage of its documents.

Thursday, April 3, 2014

What BeyondRecognition Brings to Document Management

I found this article about BeyondRecognition written by Mimi Dionne which does an excellant job of explaining in plain english how BR can benefit every business with large unstructured data collections.
You can read the entire article right here
Ever heard of BeyondRecognition? If not, the time to learn is now. The Chantilly, Va.-based "document textnology" software provider offers document managers an alternative to optical character recognition (OCR), while delivering results with accuracy and speed.

How BeyondRecognition Works




BeyondRecognition (“BR”) may be a young innovation, but it is a viable alternative to OCR. It utilizes glyphs, a letter or character formed by pixels that are of a sufficiently different color from the background of the document as to be identifiable. BR groups like glyphs into clusters at the character and word level. BR converts one glyph per cluster to text as appropriate.
While OCR continuously decides what each glyph is, BeyondRecognition’s single instance technology need only recognize one glyph per cluster to form a catalog of letters or characters. The advantage: the return on investment of using single instance recognition technology is much higher with a smaller data set — a faster processing speed and better accuracy rate — which shortens the Records Management program’s work breakdown structure significantly.
Because BeyondRecognition software is glyph dependent — not text — it is more versatile:
  • BR is language agnostic. It currently recognizes over forty languages.
  • BR is symbology agnostic. It can recognize and relate non-text elements.
  • BR clusters visual similarities. It works on all kinds of documents.
  • BR is over ninety-nine percent accurate.
  • BR scales. It can analyze millions of pages per day per the BeyondRecognition server.
BeyondRecognition’s zonal attribute extraction permits subject matter experts to extract attributes from document classifications by clicking and dragging zones on one document per document type cluster.

Again for more of Mimi's article click here

Tuesday, April 1, 2014

BeyondRecognition Denies Plans to Aquire EMC or Kofax

Independent information governance technology provider allays concerns it will seek to acquire market share through acquisitions.




Germantown TN – April 1, 2014. John Martin, CEO and Founder of BeyondRecognition, LLC, a Memphis-based technology company providing data-driven information governance technology to Fortune 500 companies, today denied trade rumors that BR had plans in place to acquire the stock or assets of either Kofax or EMC. According to Martin, “While both Kofax and EMC presently have respectable revenue numbers, we have no plans to acquire them. We believe that our organic growth will permit us to capture a significant share of their document capture and business process automation business.” 
To support his view of BR’s growth potential, Martin noted that BR had signed MSA agreements with six Fortune 100 clients in Q1, 2014.                         
Martin went on to explain that BR’s information governance technology was based on visual similarity, enabling it to automatically classify documents without the client having to develop upfront document classification rules or select multiple exemplars for each classification. “This greatly compresses the time frame required to launch projects like content migration or file share remediation. The fact that BR classifies native electronic documents as well as scanned paper documents is also a huge competitive advantage.”
About BeyondRecognition
BeyondRecognition (“BR”) provides enterprise-scale information governance technology to Fortune 500 clients. BR’s core technology classifies electronic and scanned paper documents based on their visual similarity.  Other components of BR’s offerings include zonal attribute extraction, visual deduping, and glyph recognition. BR technology enables content migration, file remediation, and other IG tasks as well as powering document-intensive business processes. BR’s clients enjoy rapid project start-up  and improved accuracy in coding or extracting document attributes, and they particularly appreciate being able to finish projects in months that had originally been scheduled to take years.
For more information about BeyondRecognition, visit the BR website at www.BeyondRecognition.net, or contact Joe Howie, VP, Corporate Communications, at jhowie@beyondreognition.net, or 918-894-6943.This release valid only on April 1, 2014 – think about it and have a great day.
You can also follow BR on Twitter @BeyondRecog or join the BeyondRecognition group on LinkedIn atwww.linkedin.com/company/beyondrecognition.
Credits: Globe in graphic obtained under Creative Commons license from all-free-download.com, "Modern Globe Blue and Green Connection Vector Illustration.jpg"

Sunday, September 2, 2012

Willie Wonka, Big Data, Pure Imagination


Hold your breath, Make a wish, Count to Three

Come with me …And you'll be
In a world of Pure imagination
Take a look, And you'll see Into your imagination

We'll begin
With a spin…Traveling in
The world of my creation…What we'll see
Will defy….Explanation





Maybe you thought that BeyondRecognition was merely the greates automatic “Visual Document Clustering” New Tool for Big Data. Well you'd be correct yet wrong. BeyondRecognition's powerful graphical engine can provide never before dreamed of  levels of image enhancement. 


BeyondRecognition's Visual-Similarity Clustering automatically processes and groups documents together for document boundary detection and document type classification, regardless of source and format — seamlessly processing native electronic files and scanned documents
Automatic Visual Document Clustering means that visually similar pages are gathered based on their graphical, rather than textual, content.  This avoids the errors normally encountered in extracted or generated text and leverages non-text graphical elements such as logos, form elements and other objects to greatly improve accuracy.

About BeyondRecognition

BeyondRecognition is a "textnology" company that has developed unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. Disclosure of further details is being deferred until one or more patents on the process are filed. BeyondRecognition is working with a select number of companies in the electronic discovery and document management industries. 
For more information, visit www.BeyondRecognition.net.

About Focus Data Management
FDM is the sales and marketing arm of BeyondRecognition. With offices on both coasts FDM is available to help you customize your Big Document Solutions using the power of BeyondRecognition. Contact us at 804.690.0010 or 562.822.7141

Tuesday, August 21, 2012

Did We Really Send that File to Opposing Counsel?


To say an overwhelming portion of documents involved in litigation or regulatory matters today are stored electronically is nearly becoming rote. How many documents are involved, how they're processed, in what format they're produced -- it's all old news.
With recent cases cropping up involving big-name companies, a more interesting conversation surrounds what mistakes are being made in the electronic discovery process and what happens with data before it is handed over to the other side.
What mistakes are we talking about, really? In the process of collecting, preserving, de-duplicating, filtering, culling, and reviewing ESI, opportunities for errors range from entering incorrect date ranges to inaccurately entering specific format requirements for tools used in downstream stages of the process. Not surprisingly, quality is a huge challenge for a discovery process that involves ever-growing volumes of data. The opportunity for mistakes, oversight, or simple carelessness comes from imperfect technology and the fallible nature of people who just can't guarantee 100 percent focus and attention to massive quantities of information.
Several high-visibility cases involving McDermott Will & Emery, Google, and Duane Reade, Inc. have drawn attention to quality control issues in the discovery process and highlighted how slip-ups can result in waiver of privilege and ultimately impact the outcome of the case. These cases reinforce the critical importance of quality control at the tail end of discovery before production to opposing counsel.
Inadvertent waiver of privilege is huge and costly in terms of damages paid and reputation tainted. But what are the other mistakes that law firms, e-discovery vendors and in-house counsel are making -- and how can they be readily fixed? The top 10 mistakes include:
1. Ineffective Redactions
Applying redactions can be tricky if you don't understand the process. Mistakes include failing to "burn-in" the redaction on the image, not updating or re-OCRing the text files to match, providing un-redacted native files, and failing to redact certain metadata. To remedy, re-review documents on your final production media to make sure the redactions are permanent. Compare the images along with the text. Check if the native is being produced -- and update it to reflect the redaction or remove it from the production. Inspect the metadata fields -- and ensure no privilege or confidential information is found.
2. Outdated Review Coding
A document may have been marked 'Privilege' by the review team, but if a request for that document to be produced occurred before a change in review coding, who knows about it? It's important to communicate changes, have a process to lock documents in the review system, and recheck the review coding against the final production media before sending it out the door.
3. Extra and/or Orphaned Files
Sometimes documents need to be manually removed from a production before it's sent off to opposing counsel. It's not enough to just remove the references from the load files. Have a process in place that checks all the file references in your load files and verifies those supporting files -- including images, text, and native -- can be matched against your load files. If you find any files that can't be matched, they may not belong in the production.
4. Wrong Starting Number
Your first two productions went smoothly, but things started to go wrong with the third production -- the "next bates number" for the third production was already used in the second -- and no one noticed until opposing counsel complained. Now you have documents with duplicate bates numbers, and every production from where the error started needs to be re-produced. Keeping accurate logs for tracking productions and using e-discovery software to compare bates numbers in past productions to new productions can flag duplicates or gaps before it's too late.
5. Missing Confidentiality
The production request has been made to the litigation support team -- and some documents have been designated 'C' for Confidential and 'HC' for Highly Confidential. Transferring the confidential coding and branding it on the images is another step in the process. Unfortunately, the team 'forgot' to do the branding step. By spot-checking the final production media -- and comparing the coding field that contains the confidentiality status against the production image -- the crisis could be adverted.
6. Wrong Production Specifications
Is everyone on the same page regarding the production specifications? Is it supposed to be single-page or multi-page images; TIFF or PDF; native files for all or just the spreadsheets; extracted text along with OCR; metadata fields included in the appropriate load files; and correct load files? If final productions meet the agreed upon production format, you will save time, money and embarrassment by not having to fix mistakes.
7. Not Producing Enough
If documents are being reviewed natively and converted to image, such as TIFF or PDF files, ensure all that can be produced is being produced. For example, spreadsheets may include hidden rows, columns, or sheets, or print areas might be set to exclude data. PowerPoint files may contain notes, or CAD files may have missing supporting files. It's possible the review team did not review everything until a document has been imaged -- or even missed relevant data that should have been deemed responsive. Your process should check the results of native files after they are converted to image.
8. Load Files Poorly Constructed 
Transferring documents from one e-discovery platform to another generally involves the use of load files. Load files move a document's file references -- including any image, text, and/or native files -- along with metadata. But with no single format commonly used or requested, knowing the intricacies of each load file requires experience. Simulating and testing the loading process -- whether it's using the target platform or a tool designed for reading load files -- can quickly identify these mistakes.
9. Failure to Communicate
Communication among all groups involved can be paramount to an inadvertent production of privileged documents. For example, the final production search may include privilege documents that need to be numbered inline and then removed from the export. But if no one checks to ensure these documents were removed after being numbered, you may end up producing privilege. The team requesting the production should always check the contents of the final production to ensure the correct number of documents and search filters have been applied before sending to opposing counsel.
10. Assuming it's Right
Quality control measures should be in place among all the teams handling a document production, whether it's the operator running the production or the attorney signing off on the production before it's sent to opposing counsel. Assuming a production doesn't inadvertently contain privilege documents -- when you had the opportunity to inspect it beforehand -- can be a costly mistake. Utilizing quality control tools designed for e-discovery, especially for productions, can allow litigation support, attorneys and paralegals to effectively and efficiently check their productions.

Justin Blessing is the Director of Product Development for Compiled Services, makers of e-discovery software ReadySuite. He can be reached at Justin@compiledservices.com. More information can be found at www.compiledservices.com. The company can be followed on Twitter @CompiledSrvcs or on Facebook at www.facebook.com/CompiledServices.

Tuesday, July 31, 2012

Dr. Stephen V. Rice Joins BeyondRecognition Advisory Board





Germantown, TN (July31, 21012). John Martin, founder and CEO of BeyondRecognition, LLC, announced today that Stephen V. Rice, Ph.D., has agreed to join BeyondRecognition’s Advisory Board and to consult with BeyondRecognition.  Martin indicated that, “One of the major functionalities provided by BeyondRecognition is our ability to extract textual and other non-textual glyphs from document images.  Dr. Rice’s scientific background uniquely qualifies him to help us to create objective measures of the accuracy of that conversion process at the character, word, and significant-word level.”

For five years, Dr. Rice conducted the first large-scale independent evaluations of commercial optical character recognition (OCR) systems while at the Information Science Research Institute of the University of Nevada, Las Vegas (UNLV). As part of that work he developed sequence comparison algorithms to measure OCR accuracy. He is author of the classic book, “Optical Character Recognition: An Illustrated Guide to the Frontier,” which is essential reading for anyone involved in developing OCR or CAPTCHA systems.

Rice noted that, “BeyondRecognition has taken a fresh approach to the challenge of extracting text from document images. They have developed several innovative technologies for document conversion and retrieval.  I look forward to assisting them as they continue to break new ground in this area.”

Martin continued, “For the past year we have been building out our code base and our infrastructure and will be looking to Dr. Rice to help us develop statistically-sound performance measures to evaluate our performance. For example, next month we plan to benchmark our system on the approximately 30 million page images of tobacco litigation documents obtained from the Text Retrieval Conference (TREC) sponsored by the National Institute of Standards and Technology (NIST). Our goal is to perform the text conversion on the 30 million pages, globally edit the text, index it and be able to perform sub-second retrieval on any page in the collection within 72 hours, start to finish. We want valid, reliable metrics to use to report the results.”

As to the significance of the document collection, Martin observed, “The TREC Tobacco documents have been used by the TREC Legal Track to gauge the efficacy of various text retrieval systems and methodologies, despite known issues of using inaccurate OCR. The studies compared the results from queries or processing to the documents identified by manual reviews, and the results have been used to argue in favor of ‘predictive coding’ or ‘technology-assisted review.’ We believe that the new, more accurate text output by BeyondRecognition will let researchers show how the early studies may have understated the effectiveness of technology-assisted review.”

About Dr. Rice

Dr. Stephen V. Rice consults as a computer scientist and software engineer with expertise in algorithms, computer audio, computer simulation, database systems, pattern recognition, programming languages, and related areas. He was a computer science professor at the University of Mississippi, was chief software engineer at the UNLV Information Science Research Institute, and is Founder and CTO of Comparisonics Corporation. He served on the board of advisors of the Federal Intelligent Document Understanding Laboratory of the U.S. Central Intelligence Agency. 

For more about Dr. Rice, see http://www.stephenvrice.com

About BeyondRecognition

BeyondRecognition has developed unique character, word, and document attribute recognition and extraction capabilities for analyzing image-based documents. Its glyph clustering and cataloging approach enables rapid, globally-editable text recognition with accuracy rates far beyond traditional OCR. BeyondRecogntion also clusters documents based on visual similarity and permits location-based, cluster-specific data element extraction for coding or abstracting data elements from the documents. Clustering by document type permits prioritized data element extraction using the powerful graphical user interface to highlight zones, and to write and instantly test and verify extraction rules.

Although nominally a “startup,” the principal technologists at BeyondRecognition have been working in the fields of document conversion, electronic evidence forensics and processing for decades. CEO John Martin was previously a founder of Cricket Technologies and RedFile LLC.
For more information, visit www.BeyondRecognition.net