Showing posts with label TAR. Show all posts
Showing posts with label TAR. Show all posts

Thursday, May 8, 2014

"Predictive" Coding and the Naked Emperor

I've always been suspicious of the claims of the "predictive coding" zealots. Every year it seems that there is a new buzzword in the field of  Ediscovery. Technology Assisted Review - Check! Information Governance - Check! Cloud Computing - Check. Big Data - Check.  

My friend John Martin did a magnificent job of distilling some of my concerns in his blog entry that can be found right here.

The Emperor has No Clothes - and PC Can't See Image-Only Documents

There are several parallels between predictive coding (AKA technology assisted review) and Hans Christian Andersons' tale, "The Emperor's New Clothes." In the story, two weavers tell the emperor they will make him a suit of clothes that will be invisible to those people who are unfit for their position, stupid, or incompetent. None of the emperor's subjects want to admit to those deficiencies so the emperor parades around with no clothes on until a child states the obvious - the emperor has no clothes.
Predictive Coding - do not see any evidence hereIn the case of predictive coding, its advocates have touted the efficacy of their approaches in white papers, blogs, and social media postings, and have practically created a separate industry to host conferences promoting the wonders of predictive coding. Few people want to ruin the moment or buck the trend by pointing out what is obvious when one considers the technology underlying predictive coding - it is completely dependent on having text to analyze.  It will absolutely fail to analyze documents for which there is no text, and will do a miserable job where the text is of poor quality.
This might be just an esoteric debating point if virtually all documents had associated text. However, in some industries like oil & gas, half or more of some collections will be engineering drawings and schematics that were output to image-only PDF for distribution and use by those who don't have the software licenses needed to view the documents in their original file formats.
Predictive Coding 100 percent right 20 percent of the timeIn practically all industries it is common practice to develop documents in one application like Word and then, once finalized, distribute them as image-only PDF so they can be viewed on a variety of devices and so recipients can't easily change the content. In one collection we analyzed, only 20% of the PDFs had associated text. Even if predictive coding were 100% effective, the most it could classify would be 20% because it literally cannot "see" the 80% without text. If in fact predictive coding has a recall rate of 70-80% of what it can see, that would mean that predictive coding would have identified 14 to 16% of the total PDFs (70% x 20% = 14% or 80% of 20% = 16%). By contrast, BR's visual classification technology classified 100% of them.
PDFs will potentially be among the most relevant file types in a collection because that is the format used to distribute information within and among groups of people within an organization, and among organizations. Note that even if in some unique e-discovery settings predictive coding is acceptable, the text-restriction failing of predictive coding will be fatal for broader information governance purposes.
So... if you're going to use predictive coding, at the very least measure what PC doesn't "see." If you're planning on using PC for information governance purposes, make sure that the organization doesn't mind not classifying a potentially significant percentage of its documents.

Thursday, April 3, 2014

Information Governance Lessons from the Six Blind Men and the Elephant


Information Governance Lessons from the Six Blind Men and the Elephant

Posted by Rich E. Davis
Mar 30, 2014 11:35:00 PM

Elephant_wBlindment_Fx600Most of us have heard about the parable of the six blind men and the elephant - it may actually be the first recorded instance of faceted classification. Six blind men touched different parts of an elephant and each described a completely different thing based on their own perspective or “view” of the elephant: The one who felt a tusk reported it as a pipe, the one who felt an ear reported a fan, the belly was reported as a wall, the trunk as a branch, a leg was reported as a pillar and the tail was described as being a rope.
This story illustrates several important information governance lessons:
Elephant_Separate_Views_Fx600Different Stakeholders Have Different Views & Needs. People’s views of and information needs from any given corpus of documents will vary according to where they are in an organization and the functions they perform. People with different roles in a company will naturally be interested in different attributes of the documents in the corpus and may well use different descriptors when describing or trying to find some of them. While some document attributes are common to all stakeholders, others, namely those which enable an individual or group to perform their job function within an organization, are probably not.
Here is an example of how different roles will be interested in different attributes:
The Offshore Power Plant
  • Stakeholders in the tax department need to know whether an expenditure on a sub-sea turbine, a critical component on a key project, can be categorized as an operating or capital expense in the jurisdiction where the project’s work is being performed.
  • The plant maintenance department needs to know:
    • When the warranty period kicks in.
    • Part numbers, nomenclature and service level.
  • Environmental Safety & Health wants to know that the team and all the contractors associated with the project are properly qualified and sanctioned to install the turbine to the engineering design specifications and operated within the tolerances.
  • RIM and Compliance wants to track the locations of all relevant project documents for information lifecycle management, disposition and regulatory reasons.
  • IT needs to ensure that business critical documents are backed up for disaster recovery and business continuity purposes.
  • InfoSec wants to know that the IT group has the requisite information to ensure that the IP associated with the project is properly secured and that the people who access the content have the proper authorization to do so.
     
Some or all of the information required by the stakeholders above will be objectively evident on the face of individual documents. Other “subjective” attributes may have to be assigned (e.g., “project lead engineer”) by knowledge workers with specific domain expertise, and other more granular data elements (e.g., installation location) may have to be assigned by linking attributes from other authoritative data sources or systems of record.
Elephant_Reconstituted_Fx600Just preserving documents without having a systematic, dynamically updatable and holistic view created by assimilating other interrelated data points will result in an incomplete picture of a project or process. Without a holistic way of assembling and viewing all the extracted document attributes of interest to the various stakeholders, the overarching information governance needs of the organization will never be met. There will be incomplete, ambiguous, erroneous and superfluous data points.
Limited Data Points Means Incomplete or Distorted Pictures. As the elephant parable illustrates, having only one or a few attributes available results in having an incomplete or distorted picture of what is being managed. The blind men’s picture is so distorted in fact, that when word of an elephant rumbling through cane fields destroying them in search of food reaches their ears, they have no adequate description for the sum of the parts, and thus no way of applying the individually assimilated knowledge in a holistic fashion. The more uniform, accurate and persistent the document attributes or facets that are available, the greater the ability of the organization to assimilate seemingly disparate information to form a more accurate picture of present and future state projects.
Elephant_Multiplied_Fx600Duplicated data sets. Without a holistic enterprise content plan, each stakeholder starts keeping their own copies of documents so they can extract the attributes they are interested in. The result is multiple copies of the same documents, multiple expenditures to extract the same attributes, and inconsistencies in ways that the same data is extracted and stored.

THE SOLUTION

The challenge described above is endemic. It exists across all types of businesses in every jurisdiction. Corporations of all sizes are dealing with big data symptoms and are stymied when comes to finding a cure that has not been available from prior technology.
Standing apart from the herd is Continuum Advisors. At Continuum, we believe in using powerful emerging technology to help our clients address their most daunting data management challenges. To that end, we have incorporated BeyondRecognition (“BR”) in our services matrix for IG, legal, information security, RIM and a host initiatives that required powerful, scalable data analytics.
BR is a radically new, data-driven information governance technology that meets the IG needs of multiple stakeholders in any enterprise, public or private. Continuum has implemented BR technology at multiple Fortune 500 clients with great success.
Elephant_BR_Consclustion_Fx600The highly experienced CA team chose to align with BR as it is the only technology in the world that automatically classifies electronic files or scanned paper documents based on their visual characteristics – and without having to waste time writing rules to identify each type of document or designating exemplars for each document type. This is tremendously important because accurate, consistent classification is the bedrock upon which all IG initiatives are built. BR solves this long-standing, previously intractable problem.
Subject matter experts can quickly determine how to classify all the documents in a document cluster by examining one or two documents per cluster. They can also associate a document type name with the cluster based on their organization’s document classification tree, and assign retention periods based on the classification.
Our subject matter experts in energy, financial services, and pharmaceuticals work with corporate knowledge workers to extract multiple attributes from each document classification by “painting,” i.e., clicking and dragging boxes, on an image of a document from each cluster. BR then automatically extracts the specified attributes and associates each attribute with the attribute or field names assigned by the subject matter experts. The extracted data can then be loaded into the appropriate content management system.
The various attributes enable the BR-processed documents to be associated with management control systems, e.g., pipeline planning and maintenance, or capital asset acquisition, or ESH inspections. The various attributes serve to provide multiple views into the document collection.
The extracted attribute values can be normalized prior to loading into the target system or the extracted values can be used to update and validate existing field authority lists.
For more information, please contact Rich E. Davis.

Tuesday, April 1, 2014

BeyondRecognition Denies Plans to Aquire EMC or Kofax

Independent information governance technology provider allays concerns it will seek to acquire market share through acquisitions.




Germantown TN – April 1, 2014. John Martin, CEO and Founder of BeyondRecognition, LLC, a Memphis-based technology company providing data-driven information governance technology to Fortune 500 companies, today denied trade rumors that BR had plans in place to acquire the stock or assets of either Kofax or EMC. According to Martin, “While both Kofax and EMC presently have respectable revenue numbers, we have no plans to acquire them. We believe that our organic growth will permit us to capture a significant share of their document capture and business process automation business.” 
To support his view of BR’s growth potential, Martin noted that BR had signed MSA agreements with six Fortune 100 clients in Q1, 2014.                         
Martin went on to explain that BR’s information governance technology was based on visual similarity, enabling it to automatically classify documents without the client having to develop upfront document classification rules or select multiple exemplars for each classification. “This greatly compresses the time frame required to launch projects like content migration or file share remediation. The fact that BR classifies native electronic documents as well as scanned paper documents is also a huge competitive advantage.”
About BeyondRecognition
BeyondRecognition (“BR”) provides enterprise-scale information governance technology to Fortune 500 clients. BR’s core technology classifies electronic and scanned paper documents based on their visual similarity.  Other components of BR’s offerings include zonal attribute extraction, visual deduping, and glyph recognition. BR technology enables content migration, file remediation, and other IG tasks as well as powering document-intensive business processes. BR’s clients enjoy rapid project start-up  and improved accuracy in coding or extracting document attributes, and they particularly appreciate being able to finish projects in months that had originally been scheduled to take years.
For more information about BeyondRecognition, visit the BR website at www.BeyondRecognition.net, or contact Joe Howie, VP, Corporate Communications, at jhowie@beyondreognition.net, or 918-894-6943.This release valid only on April 1, 2014 – think about it and have a great day.
You can also follow BR on Twitter @BeyondRecog or join the BeyondRecognition group on LinkedIn atwww.linkedin.com/company/beyondrecognition.
Credits: Globe in graphic obtained under Creative Commons license from all-free-download.com, "Modern Globe Blue and Green Connection Vector Illustration.jpg"

Wednesday, June 13, 2012

MAY BRINGS RICH C-LEVEL EXPERIENCE IN INDUSTRY AND PHILANTHROPY TO HIGH-TECH STARTUP BEYONDRECOGNITION


Germantown, TN: (May 30, 2012). John Martin, founder and CEO of BeyondRecognition, LLC, today stated that, ”BeyondRecognition is pleased to announce that Ken May will be providing business development guidance for BeyondRecognition as it pushes its innovative image-based document analysis technology into key markets like mortgage and loan processing, and the oil and gas industries.” BeyondRecognition’s breakthrough integrated workflow enables companies to obtain actionable intelligence from image-based and electronic format documents at a fraction of the cost associated with manually reviewing and abstracting paper files and often with higher accuracy and reliability.

Martin continued, “BeyondRecognition’s core competencies lie in document processing and analysis, and Ken brings an incredible wealth of experience managing FedEx Kinkos, one of the largest and most wide-spread document copying and handling operations in the world, as well as planning and managing some of the most highly automated decision-support systems in the world. He also has a wealth of C-level contacts at companies across America from his many years of service as Chairman of the National Board of Trustees for the March of Dimes. We look forward to being able to capitalize on his rich experience, energy, and industry knowledge.”

Ken May commented, “I have had the opportunity over the years to review many exciting technologies at all sorts of start-ups and emerging market leaders, but I was especially struck at how innovative BeyondRecognition’s technology is and at the incredible value it offers companies that are faced with needing to analyze and process large volumes of paper-based records. This is particularly true in industries like home loan processing where the documents in the underlying files are typically not all or even mostly electronic. The need to process existing back files of loan documents and to eventually automate the new loan initiation process represents an enormous potential. I look forward to helping spread the message about this important new technology.”

About Ken May

Beginning as a manager of hub operations for FedEx in 1982, May served in various management positions, becoming VP, Global Operations Scheduling and Control in 1996. He then served as Sr. VP, Air-Ground and Freight Services 1997 to 1999, was Sr. VP US Operations from 1999 to 2004, COO at FedEx-Kinko’s Office and Print Centers from 2004 to 2006, and was President and CEO at FedEx-Kinko’s Office and Print Centers from 2006 to 2008. 
May served as Chairman of the National Board of Trustees at the March of Dimes from 2007 to 2011, and was President of ES3, LLC, the third-party logistics subsidiary of C&S Wholesale Grocers, the eighth-largest privately held company in the US by revenue from 2010 to 2011. From 2011 to 2012 he was President and COO at Krispy Kreme Doughnuts.

May has been a director of PF Chang’s China Bistro since May 2007, and serves on the Board of Directors of Greystone Medical Group. 

For more about Ken May, see http://en.wikipedia.org/wiki/Ken_may.

About BeyondRecognition

BeyondRecognition has developed unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. Its glyph clustering and cataloging approach enables rapid, globally-editable text recognition with accuracy rates far beyond traditional OCR. BeyondRecognition also clusters documents based on visual similarity and permits location-based, cluster-specific data element extraction for coding or abstracting data elements from the documents. Clustering by document type permits prioritized data element extraction using the powerful graphical user interface to highlight zones, and to write and instantly test and verify extraction rules.

Although nominally a “startup,” the principal technologists at BeyondRecognition have been working in the fields of document conversion, electronic evidence forensics and processing for decades. CEO John Martin was a founder of Cricket Technologies, LLC and RedFile LLC.

For more information, visit www.BeyondRecognition.net

Wednesday, June 6, 2012

Unlocking Paper Based Intelligence with Disruptive Technology

"You want your documents to be searchable - not laughably searchable . . . " John Martin



When confronted with massive amounts of unstructured data and the need to access the business intelligence locked inside that data the options before today were expensive, required massive human intervention, were extremely time consuming and most troubling, very ineffective. John Martin loves disruptive technology. I love how John Martin thinks.

John's latest game changing software BeyondRecognition is being referred to as a "Big Data innovator" by several of the " Big Four" accounting  firms. The tool was originally built  for a  company that  needed to extract key information from a 30+ year old scanned paper document set of 2.3 BILLION pages for a due diligence effort.  In the energy sector, Beyond Recognition's  glyph clustering technology makes it possible to search for symbols 
used on maps to indicate things like radioactive wells, salt-water wells or API number codes.

As a result of this disruptive new technology Martin notes , "We're seeing a great deal of interest in this approach in the mortgage and energy sectors. The mortgage industry in particular typically has a relatively finite number of documents in loan files supporting the loan decisions, with definable types of data being of interest on each type of document. Our process could greatly lower the cost of tracking all those data elements during loan initiation, or to quality control the file for audit or sale purposes."




Essentially BeyondRecognition's  unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. In Plain english  BR allows clients to extract very valuable business intelligence from scanned and digitized files fast, accurately and in a cost-efficient manner. In a press release Barbara Johnson, CFA, former executive of USAA Federal Savings Bank, serving as Chief Credit Officer and Senior Vice President of Real Estate Lending Services and now a Principal with Saccadent, a Financial Consulting firm, has reviewed the clustering and data extraction capabilities and offered the comment that, "In today's environment the ability to extract, utilize and match data across a variety of documents is incredibly powerful. This type of technology offers the promise of significantly decreasing the time and cost to process a thoroughly compliant loan from application through origination, audit, sale and servicing. An automated system to confirm all the critical items match throughout the process and are in the appropriate format and location on all documents would be invaluable. Lending is a document-rich industry and the time is perfect for this type of technology."

Another intriguing aspect is that  BR solution is language agnostic automatically recognizing 40+ languages interspersed throughout any data set with no up-front programming required and performs at 99.5% word accuracy on first pass unassisted review. The solution can scale to meet customer requirements between 500k - 50m pages per day regardless of the legibility of the images. In fact using BR Adaptive Image Enhancement techniques restorion of  poor quality document images to much improved legibility is a seemless byproduct 


For additional information or to request a demo  contact John Martin at John AT beyond recognition DOT net
or Michael Mulcahy at Michael@focusdata-mgt.com or by phone at 562 546-2465