Showing posts with label Technology Assisted Review. Show all posts
Showing posts with label Technology Assisted Review. Show all posts

Monday, June 2, 2014

Superior Document Services announces two team members became Relativity Certified Trainers

Superior Documents Services is proud to announce that two members of our team, Daniel Attaway and Bryant Overgard, have recently become Relativity Certified Administrators! The Relativity Certified Administrator (RCA) program ensures that case administrators fully understand Relativity's capabilities, allowing you to maximize the software's flexibility and provide an intuitive interface for end users. To obtain certification, you must earn a score of 80% or higher on the RCA exam, which contains both written and practical elements
By completing the Relativity Certified Administrator (RCA) program, these individuals have shown a thorough understanding of Relativity and how to best leverage the platform to our clients’ advantage. This is just part of our ongoing goal to provide truly Superior service to all of our clients.




Thursday, May 8, 2014

"Predictive" Coding and the Naked Emperor

I've always been suspicious of the claims of the "predictive coding" zealots. Every year it seems that there is a new buzzword in the field of  Ediscovery. Technology Assisted Review - Check! Information Governance - Check! Cloud Computing - Check. Big Data - Check.  

My friend John Martin did a magnificent job of distilling some of my concerns in his blog entry that can be found right here.

The Emperor has No Clothes - and PC Can't See Image-Only Documents

There are several parallels between predictive coding (AKA technology assisted review) and Hans Christian Andersons' tale, "The Emperor's New Clothes." In the story, two weavers tell the emperor they will make him a suit of clothes that will be invisible to those people who are unfit for their position, stupid, or incompetent. None of the emperor's subjects want to admit to those deficiencies so the emperor parades around with no clothes on until a child states the obvious - the emperor has no clothes.
Predictive Coding - do not see any evidence hereIn the case of predictive coding, its advocates have touted the efficacy of their approaches in white papers, blogs, and social media postings, and have practically created a separate industry to host conferences promoting the wonders of predictive coding. Few people want to ruin the moment or buck the trend by pointing out what is obvious when one considers the technology underlying predictive coding - it is completely dependent on having text to analyze.  It will absolutely fail to analyze documents for which there is no text, and will do a miserable job where the text is of poor quality.
This might be just an esoteric debating point if virtually all documents had associated text. However, in some industries like oil & gas, half or more of some collections will be engineering drawings and schematics that were output to image-only PDF for distribution and use by those who don't have the software licenses needed to view the documents in their original file formats.
Predictive Coding 100 percent right 20 percent of the timeIn practically all industries it is common practice to develop documents in one application like Word and then, once finalized, distribute them as image-only PDF so they can be viewed on a variety of devices and so recipients can't easily change the content. In one collection we analyzed, only 20% of the PDFs had associated text. Even if predictive coding were 100% effective, the most it could classify would be 20% because it literally cannot "see" the 80% without text. If in fact predictive coding has a recall rate of 70-80% of what it can see, that would mean that predictive coding would have identified 14 to 16% of the total PDFs (70% x 20% = 14% or 80% of 20% = 16%). By contrast, BR's visual classification technology classified 100% of them.
PDFs will potentially be among the most relevant file types in a collection because that is the format used to distribute information within and among groups of people within an organization, and among organizations. Note that even if in some unique e-discovery settings predictive coding is acceptable, the text-restriction failing of predictive coding will be fatal for broader information governance purposes.
So... if you're going to use predictive coding, at the very least measure what PC doesn't "see." If you're planning on using PC for information governance purposes, make sure that the organization doesn't mind not classifying a potentially significant percentage of its documents.

Thursday, April 3, 2014

What BeyondRecognition Brings to Document Management

I found this article about BeyondRecognition written by Mimi Dionne which does an excellant job of explaining in plain english how BR can benefit every business with large unstructured data collections.
You can read the entire article right here
Ever heard of BeyondRecognition? If not, the time to learn is now. The Chantilly, Va.-based "document textnology" software provider offers document managers an alternative to optical character recognition (OCR), while delivering results with accuracy and speed.

How BeyondRecognition Works




BeyondRecognition (“BR”) may be a young innovation, but it is a viable alternative to OCR. It utilizes glyphs, a letter or character formed by pixels that are of a sufficiently different color from the background of the document as to be identifiable. BR groups like glyphs into clusters at the character and word level. BR converts one glyph per cluster to text as appropriate.
While OCR continuously decides what each glyph is, BeyondRecognition’s single instance technology need only recognize one glyph per cluster to form a catalog of letters or characters. The advantage: the return on investment of using single instance recognition technology is much higher with a smaller data set — a faster processing speed and better accuracy rate — which shortens the Records Management program’s work breakdown structure significantly.
Because BeyondRecognition software is glyph dependent — not text — it is more versatile:
  • BR is language agnostic. It currently recognizes over forty languages.
  • BR is symbology agnostic. It can recognize and relate non-text elements.
  • BR clusters visual similarities. It works on all kinds of documents.
  • BR is over ninety-nine percent accurate.
  • BR scales. It can analyze millions of pages per day per the BeyondRecognition server.
BeyondRecognition’s zonal attribute extraction permits subject matter experts to extract attributes from document classifications by clicking and dragging zones on one document per document type cluster.

Again for more of Mimi's article click here

Tuesday, April 1, 2014

BeyondRecognition Denies Plans to Aquire EMC or Kofax

Independent information governance technology provider allays concerns it will seek to acquire market share through acquisitions.




Germantown TN – April 1, 2014. John Martin, CEO and Founder of BeyondRecognition, LLC, a Memphis-based technology company providing data-driven information governance technology to Fortune 500 companies, today denied trade rumors that BR had plans in place to acquire the stock or assets of either Kofax or EMC. According to Martin, “While both Kofax and EMC presently have respectable revenue numbers, we have no plans to acquire them. We believe that our organic growth will permit us to capture a significant share of their document capture and business process automation business.” 
To support his view of BR’s growth potential, Martin noted that BR had signed MSA agreements with six Fortune 100 clients in Q1, 2014.                         
Martin went on to explain that BR’s information governance technology was based on visual similarity, enabling it to automatically classify documents without the client having to develop upfront document classification rules or select multiple exemplars for each classification. “This greatly compresses the time frame required to launch projects like content migration or file share remediation. The fact that BR classifies native electronic documents as well as scanned paper documents is also a huge competitive advantage.”
About BeyondRecognition
BeyondRecognition (“BR”) provides enterprise-scale information governance technology to Fortune 500 clients. BR’s core technology classifies electronic and scanned paper documents based on their visual similarity.  Other components of BR’s offerings include zonal attribute extraction, visual deduping, and glyph recognition. BR technology enables content migration, file remediation, and other IG tasks as well as powering document-intensive business processes. BR’s clients enjoy rapid project start-up  and improved accuracy in coding or extracting document attributes, and they particularly appreciate being able to finish projects in months that had originally been scheduled to take years.
For more information about BeyondRecognition, visit the BR website at www.BeyondRecognition.net, or contact Joe Howie, VP, Corporate Communications, at jhowie@beyondreognition.net, or 918-894-6943.This release valid only on April 1, 2014 – think about it and have a great day.
You can also follow BR on Twitter @BeyondRecog or join the BeyondRecognition group on LinkedIn atwww.linkedin.com/company/beyondrecognition.
Credits: Globe in graphic obtained under Creative Commons license from all-free-download.com, "Modern Globe Blue and Green Connection Vector Illustration.jpg"

Wednesday, June 13, 2012

MAY BRINGS RICH C-LEVEL EXPERIENCE IN INDUSTRY AND PHILANTHROPY TO HIGH-TECH STARTUP BEYONDRECOGNITION


Germantown, TN: (May 30, 2012). John Martin, founder and CEO of BeyondRecognition, LLC, today stated that, ”BeyondRecognition is pleased to announce that Ken May will be providing business development guidance for BeyondRecognition as it pushes its innovative image-based document analysis technology into key markets like mortgage and loan processing, and the oil and gas industries.” BeyondRecognition’s breakthrough integrated workflow enables companies to obtain actionable intelligence from image-based and electronic format documents at a fraction of the cost associated with manually reviewing and abstracting paper files and often with higher accuracy and reliability.

Martin continued, “BeyondRecognition’s core competencies lie in document processing and analysis, and Ken brings an incredible wealth of experience managing FedEx Kinkos, one of the largest and most wide-spread document copying and handling operations in the world, as well as planning and managing some of the most highly automated decision-support systems in the world. He also has a wealth of C-level contacts at companies across America from his many years of service as Chairman of the National Board of Trustees for the March of Dimes. We look forward to being able to capitalize on his rich experience, energy, and industry knowledge.”

Ken May commented, “I have had the opportunity over the years to review many exciting technologies at all sorts of start-ups and emerging market leaders, but I was especially struck at how innovative BeyondRecognition’s technology is and at the incredible value it offers companies that are faced with needing to analyze and process large volumes of paper-based records. This is particularly true in industries like home loan processing where the documents in the underlying files are typically not all or even mostly electronic. The need to process existing back files of loan documents and to eventually automate the new loan initiation process represents an enormous potential. I look forward to helping spread the message about this important new technology.”

About Ken May

Beginning as a manager of hub operations for FedEx in 1982, May served in various management positions, becoming VP, Global Operations Scheduling and Control in 1996. He then served as Sr. VP, Air-Ground and Freight Services 1997 to 1999, was Sr. VP US Operations from 1999 to 2004, COO at FedEx-Kinko’s Office and Print Centers from 2004 to 2006, and was President and CEO at FedEx-Kinko’s Office and Print Centers from 2006 to 2008. 
May served as Chairman of the National Board of Trustees at the March of Dimes from 2007 to 2011, and was President of ES3, LLC, the third-party logistics subsidiary of C&S Wholesale Grocers, the eighth-largest privately held company in the US by revenue from 2010 to 2011. From 2011 to 2012 he was President and COO at Krispy Kreme Doughnuts.

May has been a director of PF Chang’s China Bistro since May 2007, and serves on the Board of Directors of Greystone Medical Group. 

For more about Ken May, see http://en.wikipedia.org/wiki/Ken_may.

About BeyondRecognition

BeyondRecognition has developed unique character, word and document attribute recognition and extraction capabilities for analyzing image-based documents. Its glyph clustering and cataloging approach enables rapid, globally-editable text recognition with accuracy rates far beyond traditional OCR. BeyondRecognition also clusters documents based on visual similarity and permits location-based, cluster-specific data element extraction for coding or abstracting data elements from the documents. Clustering by document type permits prioritized data element extraction using the powerful graphical user interface to highlight zones, and to write and instantly test and verify extraction rules.

Although nominally a “startup,” the principal technologists at BeyondRecognition have been working in the fields of document conversion, electronic evidence forensics and processing for decades. CEO John Martin was a founder of Cricket Technologies, LLC and RedFile LLC.

For more information, visit www.BeyondRecognition.net. 

Wednesday, May 16, 2012

Document Clustering Using Facial Recognition Principles & Post-Clustering Data Extraction By John Martin

…the new document textnology

John Martin turns the Technology Assisted Review, Post Clustering Data Extraction Upside Down using Facial Recognition Principles :








Companies achieve many benefits when theycan quickly, accurately and inexpensively cluster like documents. For example, 500-page home loan files can be scanned on high speedscanners, the pages programmatically grouped intodocuments, and like document types clustered or categorized across all the loan files without significant operator intervention, permitting the company to check if the mortgage files contain the expected types of documents, e.g. deeds, inspection reports, insurance binders, lien releases, etc.

Furthermore, with precise clustering of the same document types, data that was unique to each
document can be extracted and compared to expected values. To continue the loan file example
the form of the grantee’s name or the property description can be compared across all documents or
against a loan tracking system to make sure all the values matched.


New technology from BeyondRecognition (“BR”) provides this type of functionality. BR examines the
“faces” or images of documents and using between roughly a thousand and fifteen hundred points of
comparison, clusters or groups those documents that have a high degree of similarity. By contrast, facial
recognition of picture of human faces typically uses less than 200 points of comparison.


One of the most subtle but profound consequences of programmatic document typing or clustering is that
documents are arranged in clusters before you have to decide what to name them or what data elements to
extract from them or how long to retain them.



Post-Clustering Data Element Extraction


Anyone who has ever tried to write coding rules or document processing instructions for the handling of
heterogeneous document collections knows that one of the largest obstacles is that you don’t know what’s
in the collection until you’ve gone through it all. The typical scenario is to sample documents, write the
best manual you can up front and then be prepare for many updates and revisions as the project unfolds.

The graphical interface used by BR for creating data extraction rules permits rapid, extremely accurate
coding of the various document clusters using cluster specific location coordinates and other textual as well
as nontextual markers, to extract those data elements most pertinent to each data type or cluster.

In fact, it will typically take an experienced clustering consultant less time to actually create and apply the
rules after the documents have been programmatically clustered than it would for a projectmanager to write a comprehensive coding manua before or during the project implementation phase for manual coding.

With the breakthrough graphical interface the clustering consultant can also see immediately what
data is extracted across an entire cluster right as clusters are being analyzed. This greatly reduces
the time required to write and debug data element extraction rules. Without the interactive feature,
testing would have to be done on a batch basis which is inherently more time consuming and permits far
fewer iterations.


Clustering also readily identifies anomalousdocuments, ones that are truly unique within the
collection and for which individual rules need not be developed. Anyone who has had to work within the limitations of writing data extraction rules for documents based solely upon the textual representation of those documents will be pleased to know that BR’s extraction rules can be based on non-textual graphical
elements, e.g. a logos or vertical lines. This ability to use non-text graphical elements or data
extraction flags or cues greatly expands the flexibility and power of BR’s data element extraction capability.
It enables BR to achieve an accuracy and thoroughness rivaling if not in many cases exceeding
that of human-based reviewers or coders – and at a fraction of the cost and time, to say nothing of
security benefits of not having legions of people reviewing potentially sensitive corporate documents
and trade secrets just to have indices prepared.


If you are not familiar with this technology, please visit http://www.beyondrecognition.net/.

Alternatively, to learn how you can put BeyondRecognition to work solving your document problems

Contact John Martin at John@beyondrecognition.net or Kriss@focusdata-mgt.com