- Nationwide Digital Forensic & Cyber Investigation Services
Collection produces raw data. Processing produces evidence a review team can actually work with. Elite Digital Forensics ingests collected material, expands containers, removes system files and duplicates, extracts text and metadata, runs optical character recognition, threads conversations, normalizes time zones and reports exceptions honestly, so volume drops and defensibility rises.
Updated August 2026 · Reviewed by Elite Digital Forensics examiners · Remote and on site service nationwide
Quick answer. E-Discovery processing is the stage between collection and review. Collected data is ingested with hash verification, archives and mail containers are expanded, known system files are removed using the National Software Reference Library hash set, duplicates are removed within and across custodians, text and metadata are extracted, images and scanned documents are put through optical character recognition, email is threaded, timestamps are normalized to a single stated time zone and encrypted or corrupt items are reported as exceptions rather than silently dropped. Every reduction step is quantified so counsel can state, on the record, how a volume went from raw collection to review population.
| Question | Short answer |
|---|---|
| What is DeNISTing? | Removing known operating system and application files using the NIST reference hash set. |
| What is deduplication? | Removing identical items by hash, either within a custodian or across the whole population. |
| Why thread email? | Threading groups a conversation so reviewers read once instead of reading every forward. |
| Why does OCR matter? | Scanned documents and images contain no searchable text until optical recognition adds it. |
| What are exceptions? | Items that cannot be processed, typically encrypted, corrupt or an unsupported format. |
| Which time zone is used? | One agreed zone, stated in the production documentation, usually coordinated universal or case local. |
| Does processing change evidence? | No. The forensic source is preserved and processing works from verified copies. |
| Can volume be reduced further? | Yes, with date ranges, custodian limits, file type filters and agreed search terms. |
| Step | What happens | Why it matters |
|---|---|---|
| Ingestion and verification | Hash values are checked against the collection manifest and item counts reconciled | Establishes that the processed set matches what was collected |
| Container expansion | ZIP, RAR, PST, OST, MBOX and nested archives are expanded recursively | Responsive documents commonly sit several containers deep |
| Family preservation | Parent and child relationships between messages and attachments are retained | Producing an attachment without its parent invites a completeness challenge |
| DeNISTing | Known operating system and application files are removed by hash | Removes large volumes of noise that no reviewer needs to see |
| Deduplication | Identical items are removed by hash, within or across custodians | Cuts review cost, provided custodian tracking is preserved |
| Text and metadata extraction | Body text and metadata fields are extracted into the review index | Enables search, filtering and metadata based analysis |
| Optical character recognition | Scanned documents and images are converted to searchable text | Without it, image based documents are invisible to search terms |
| Email threading | Messages are grouped into conversations with inclusive message identification | Reviewers read a conversation once rather than every reply and forward |
| Time zone normalization | Timestamps are converted to a single stated zone | Prevents chronologies that appear inconsistent across custodians |
| Exception reporting | Unprocessable items are logged with reason and disposition | Preserves defensibility and tells counsel what still needs attention |
| Search term application | Agreed terms and filters are applied with hit counts by term | Supports negotiation and demonstrates proportionality |
| Review set delivery | Output is delivered in the platform's expected load file format | Review starts without rework |
Review is the most expensive phase of discovery, and it is priced by volume. Processing is where volume is legitimately reduced, and each reduction technique needs to be defensible and quantified.
DeNISTing eliminates operating system and application files that were never authored by a custodian.
Identical items are removed by hash, with custodian membership tracked so no custodian's possession is lost.
Documents that differ slightly are grouped so review is consistent across versions.
Inclusive message identification allows a reviewer to read the final message rather than each forward.
Agreed ranges and custodian lists remove data that falls outside the matter.
Terms, file type filters and domain filters reduce the population, with hit counts reported for negotiation.
We report the funnel numerically: collected volume, post expansion, post DeNIST, post deduplication, post filtering and final review population. That table is often the most useful single document counsel has when arguing proportionality.
Every processing job produces items that cannot be processed automatically. The professional response is to report them and resolve what can be resolved, not to let them disappear from the count.
The exception report identifies each item, the reason it failed, the remediation attempted and the final disposition. If an item remains unprocessed, that fact is disclosed rather than buried, which is what prevents an argument about a silently missing document later.
Metadata frequently answers the question in dispute: who authored it, when it changed, where it came from and who else received it. Processing must preserve it rather than overwrite it, which is exactly what copying files by hand tends to do.
| Category | Representative fields |
|---|---|
| From, to, copy, blind copy, subject, sent and received times, message identifier, conversation index | |
| Documents | Author, last modified by, created, modified and last accessed times, application and revision data |
| File system | Original path, file name, extension, size and hash value |
| Custodian and source | Custodian name, source device or account and collection date |
| Family | Parent identifier, attachment count and attachment order |
| Processing | Deduplication status, exception status, optical recognition status and text extraction source |
The metadata fields to be produced should be agreed in the ESI protocol before processing begins. Agreeing them afterwards commonly forces reprocessing that a short conversation would have avoided.
This page is part of the Elite Digital Forensics E-Discovery services hub. Related coverage:
We build the processing specification with counsel before we ingest anything, so metadata fields, deduplication scope, time zone and output format are settled in advance. Data is ingested with hash verification, culled with quantified steps, threaded and normalized, and delivered in the load file format your review platform expects. Exceptions are reported honestly with remediation attempted, and the full funnel is documented so the reduction from collection to review is defensible on the record.
Elite Digital Forensics is an independent digital forensics firm providing nationwide E-Discovery services, computer and mobile device forensics, cloud and email investigations and expert witness testimony. Our examiners include former law enforcement forensic examiners and court qualified expert witnesses. We work for law firms on both sides of the docket, for corporations and in house legal departments, and for insurers. When retained through counsel, our work is generally treated as attorney work product prepared in anticipation of litigation.
It converts collected raw data into a searchable, reviewable population. That includes verifying hashes on ingestion, expanding archives and mail containers, removing known system files and duplicates, extracting text and metadata, running optical character recognition on images, threading email conversations, normalizing timestamps to a single stated time zone, reporting items that could not be processed and delivering output in the load file format the review platform expects.
DeNISTing removes files whose hash values match the National Software Reference Library, a published reference set of known operating system and application files. Those files are vendor supplied program components rather than custodian authored documents, so removing them is a standard and well accepted reduction step. The count removed is reported so the step is transparent.
It depends on the matter. Global deduplication across all custodians produces the smallest review population and the lowest cost, but custodian membership must be tracked so it remains possible to show which custodians possessed a document. Per custodian deduplication preserves possession more visibly at the cost of higher volume. This is a protocol decision worth making deliberately rather than by default.
Because collected data carries timestamps in different zones and formats depending on the system that recorded them. Without normalization, a single email thread can appear to contain replies that precede their originals, and a chronology across custodians becomes unreliable. We normalize to one agreed zone and state that zone in the production documentation.
They are logged as exceptions and worked, not ignored. Where the client can supply credentials, the items are processed normally. Where they cannot and the engagement authorizes it, password recovery may be attempted. Anything that remains unprocessed appears in the exception report with the reason and the remediation attempted, so its absence from the review population is disclosed.
No. The forensic source, whether a disk image, mobile extraction or cloud export, is preserved and hash verified, and processing works from verified working copies. Original hash values remain available for comparison, which is what supports authentication under Federal Rules of Evidence 901 and 902(14).
Yes, and it is a common engagement. We verify what arrived against whatever manifest and hash values exist, document the condition it arrived in, and report any gaps in the collection record. Where the prior collection is deficient, that finding is itself relevant, and it is better to identify it during processing than during a deposition.
It varies widely with data type and custodian behavior, so a promised percentage would be guesswork. What we do commit to is reporting the funnel numerically: collected volume, post expansion, post DeNIST, post deduplication, post filtering and final review population. That table lets counsel argue proportionality from evidence rather than impression.
#DigitalForensics #ComputerForensics #CellPhoneForensics #ExpertWitness #DigitalForensicExperts #EliteDigitalForensics #ForensicInvestigation #EDiscovery #EDiscoveryServices #ESI #ElectronicDiscovery #ChainOfCustody #ForensicCollection #LitigationSupport #ESIPreservation
This content is for educational and informational purposes only and does not constitute legal advice. Elite Digital Forensics provides independent digital forensic and E-Discovery services and expert witness testimony; we do not provide legal representation. Every case is fact specific; outcomes depend on the evidence, jurisdiction, and counsel. Retain qualified legal counsel for advice about your matter.
Elite Digital ForensicsΒ is a Professional Digital Forensics and Cyber Consulting Company that provides services nationwide.Β
Elite Digital Forensics Assistant
By submitting this form, you consent to be contacted by email, text, or phone. Your information is kept secure and confidential. Reply Stop to opt out at anytime.Β
IMPORTANT: Please remember to check your spam or junk folder
We use cookies for site functionality and, only with your permission, analytics and advertising. See our Privacy Policy for details. California residents have the right to Do Not Sell or Share My Personal Information.