How to Calculate Discovery Document Volume Cost: A Practitioner’s DIY Method From Scratch

The Core Formula You Need Before Opening Any Calculator

If you are asking how to calculate discovery document volume cost, the blunt answer is: you must translate document count into data size, layer on task-based unit rates, and model at least three vendor pricing structures before committing a budget. In my first major antitrust matter in 2017, I estimated $45,000 based on 200,000 documents at a flat per-document rate, only to receive a final invoice for $112,000 because I ignored processed data size and hidden OCR charges.

The defensible manual formula I now use is:

Total Cost = (Document Count ÷ Docs-per-GB Factor) × Storage/Processing Rate + Σ(Task Volume × Unit Rate) + (Review Hours × Blended Hourly Rate)

That equation looks tidy, but each variable conceals traps that ready-made calculators obscure. Across more than 30 matters—including a 2.3-million-document FCPA investigation and a multistate product liability suit with 400 custodians—I have refined a step-by-step methodology that you can build in Excel or replicate with our Discovery Document Volume Cost Calculator. This guide fills the gap left by vendor tools that hide the math.

Step 1: Convert Document Count to Data Volume (Docs → GB)

The first missing piece in most ranking articles is the mechanical conversion from ‘documents’ to ‘gigabytes.’ Vendors bill storage per-GB; corporate clients think in emails and contracts. You must bridge that units gap manually.

Average file sizes vary by source. In a 2021 employment case, I measured actual exports from Microsoft 365: plain text emails averaged 48 KB, HTML emails with logos 92 KB, Word docs 210 KB, Excel files 380 KB, and scanned PDFs 1.4 MB per page. A mixed corporate corpus typically lands at 250–400 KB per document native, but if imaged/TIFFed, it explodes to 2–5 MB per page.

Build a Weighted Average Conversion Factor

List your expected document mix with real percentages. Example from a typical internal investigation:

  • 40% email (avg 75 KB)
  • 25% Word/RTF (avg 250 KB)
  • 15% PDF native (avg 600 KB)
  • 10% spreadsheets (avg 400 KB)
  • 10% scanned images/TIFF (avg 2 MB per page, assume 3 pages/doc = 6 MB)

Weighted average KB/doc = (0.4×75)+(0.25×250)+(0.15×600)+(0.1×400)+(0.1×6000) = 30+62.5+90+40+600 = 822.5 KB. That is ~0.8 MB/doc. Thus 100,000 docs ≈ 82 GB native; but if the scanned fraction rises to 25%, the average jumps to 1.9 MB/doc and 190 GB.

Most people don’t realize that ‘duplicate’ files counted by a client as separate documents often share identical binary data; however, email attachments extracted from PST containers can multiply child items by 3–5 times. I once had a 50,000-email custodian produce 190,000 separate load-file items after attachment breaking and thread expansion.

Account for Family Grouping and Container Explosion

When processing PST, ZIP, or OST containers, a single business ‘document’ becomes many records. Your cost model must adopt the vendor’s counting method—usually ‘items’ not ‘matters.’ Ask for their definition in writing before estimating. In a 2019 matter, our counted items exceeded the client’s document estimate by 38% solely due to container expansion.

Case Study: Converting 2.3 Million Documents in an FCPA Matter

To make the conversion concrete, here is a redacted snapshot from a 2022 foreign corrupt practices investigation. The client estimated ‘about 2 million documents.’ After custodian collection we recorded 2.31 million items:

  • 1.1M emails (avg 85 KB) = 93.5 GB
  • 420k Word/PowerPoint (avg 280 KB) = 117.6 GB
  • 300k Excel (avg 420 KB) = 126 GB
  • 290k scanned PDFs (avg 4 MB/page, 2.5 pages) = 2,900 GB
  • 200k misc (avg 500 KB) = 100 GB

Raw total ≈ 3,337 GB. After processing index overhead of 18%, billable storage was 3,938 GB. The client’s original ‘document’ count ignored the scanned PDF blow-up; per-GB pricing would have been catastrophic without aggressive dedupe. This case taught me to always separate scanned page volume from native doc count.

Step 2: Factor Hidden Processing Tasks That Inflate Volume Cost

Competitor articles mention OCR and dedupe but rarely quantify them in a manual calculation. Here is the task list that actually hits budgets, with 2024 market rates I negotiated:

  • Imaging/TIFF conversion: $0.02–$0.05 per page if not native.
  • OCR of scanned PDFs: $0.01–$0.03 per page; a 6 MB scanned doc at 300 DPI is ~3 pages = $0.06–$0.09 each.
  • Exact deduplication: often bundled per-GB, but if per-doc, $0.005–$0.01/doc.
  • Near-duplicate analysis: adds 10–15% to processing cost.
  • Email threading: $0.02–$0.04 per email item.
  • Privilege review platform fees: per ‘highlight’ or per doc flagged, often $0.10–$0.25/doc on top of attorney time.
  • Redaction & slip-sheeting: $1.50–$3.00 per redacted page plus log creation hourly.

The thing nobody tells you: aggressive dedupe can cut volume 30–40% at the item level, but when you restore family groups and thread parents, net reduction is often only 15–20%. I learned this after promising a client a 40% savings that evaporated at production time.

Never apply a flat dedupe percentage to your cost baseline without modeling the family restoration factor and near-duplicate clustering.

Also, data volume for storage is typically charged on the processed dataset size, not original. Processing indices add 10–20% overhead. If you estimate 100 GB raw, budget 115–120 GB billable. Many firms miss this and understate hosting by five figures annually.

Step 3: Compare Pricing Models Side-by-Side With a Real Case

Vendors quote three primary structures: per-document, per-GB, and hourly. Hybrids exist. Below is a side-by-side using a real 350,000-document case I handled in 2022 (mix from Step 1, avg 0.82 MB/doc → ~287 GB processed with overhead).

Scenario A: Pure Per-Document

Rate $0.18/doc for processing + first-year hosting. Cost = 350,000 × 0.18 = $63,000. Attorney review billed separately.

Scenario B: Pure Per-GB

Rate $250/GB/year hosting+processing. Cost = 287 × 250 = $71,750. Plus per-doc tasks: OCR at $0.02/page; assume 30% docs scanned 3 pages = 315,000 pages × $0.02 = $6,300. Total $78,050.

Scenario C: Hourly + Base

Platform base $5,000, plus 400 hours at $145/hr (paralegal review supervision) = $58,000, plus storage $1,200. Total $64,200 but unpredictable if review scope expands.

Comparison summary:

  • Per-doc: predictable if volume fixed, punishes large scan sets.
  • Per-GB: rewards dedupe, punishes rich media and PDFs.
  • Hourly: best for uncertain scope, worst for sticker shock.

In that case, per-doc won at fixed scope; but when we later added 200k docs, per-GB would have saved 12%. As we covered in our Discovery Document Volume Cost Calculator, dynamic switching needs threshold modeling.

Step 4: Forecast Pre-Collection Volume With Custodian Sampling

The biggest missing piece in ranking articles is forecasting volume before you sign a vendor. You cannot pick per-GB vs per-doc blindly. I use a sampling method rooted in statistical confidence intervals.

Select a random 10% sample of identified custodians (minimum 3). Pull their last 12 months of mailbox sizes and shared drives via PowerShell or M365 compliance search. In a 2023 healthcare matter, 4 sampled custodians averaged 14.2 GB mailboxes; extrapolated to 42 custodians = 596 GB, but item counts suggested 880,000 docs.

Calculate Docs-per-GB for Your Environment

From sample, divide item count by GB. We got 1,477 docs/GB. That factor feeds Step 1. If your factor is <1,000 docs/GB, per-doc pricing likely cheaper; >2,000 docs/GB, per-GB wins. This single ratio has steered me away from bad contracts twice.

Uncertainty: sampling error ±25% at 10% sample. Use EDRM guidelines for proportionate scope. The RAND study on e-discovery costs found unpredictability drove 70% of budget overruns, reinforcing why forecasting matters.

Advanced Forecasting: Confidence Intervals and Custodian Tiers

Not all custodians are equal. I tier them: Tier 1 (executives) often hold 5× the data of Tier 3 (line staff). A simple random sample may miss extremes. Use stratified sampling: pick 2 from each tier. Compute mean and standard deviation per tier, then pool.

In a 2024 securities case, Tier 1 avg 38 GB, Tier 2 12 GB, Tier 3 4 GB. With 5/20/75 distribution, expected total = (5×38)+(20×12)+(75×4)=190+240+300=730 GB. Confidence interval 95% ±18%. That range informed our negotiation for a per-GB cap with burst allowance.

The takeaway: pre-collection forecasting is not guesswork; it is applied statistics. Spreadsheets can compute STDEV and CONFIDENCE.NORM easily.

Step 5: Build Your Own Excel Template (Or Use Our Tool)

You do not need a black-box calculator. Create a workbook with three tabs: Inputs, Conversion, Scenarios. On Inputs, list doc mix percentages and average KB. On Conversion, compute weighted KB and total GB with overhead multiplier. On Scenarios, replicate Step 3 formulas with live cells.

I have shared this template with dozens of firm CFOs. One tip: use Excel’s DATA TABLE to toggle pricing models. If you prefer not to maintain formulas, our Discovery Document Volume Cost Calculator mirrors this logic with preset rates from 2024 market surveys.

Key Cells to Include

  • Raw doc count by type (not just total).
  • Scanned page multiple (default 3).
  • Dedupe net factor (default 0.8 meaning 20% removal).
  • Privilege review hourly blended rate and platform per-doc fee.
  • Vendor per-GB vs per-doc toggle with threshold field.

Production Format Trade-offs: Native, TIFF, PDF

Many cost models stop at processing. But production format can swing cost 20–30%. Native production is cheapest but may expose metadata; TIFF+LOADFILE requires imaging every page. In a 2020 matter, converting 150,000 docs to TIFF added $22,500 at $0.015/page avg 3 pages. PDF with text layer falls between.

Consider the requesting party’s expectations and court rules. Some jurisdictions default to TIFF. Factor this into the per-document rate if vendor bundles imaging. I always add a line item ‘production format conversion’ even if zero, to force the conversation.

Using FRCP Proportionality to Challenge or Defend Costs

Under Federal Rule of Civil Procedure 26(b)(2)(B), parties need not produce electronically stored information if the burden or cost outweighs its likely benefit. That rule is your leverage when vendor quotes balloon. I have used a documented volume-cost calculation to convince a magistrate to limit custodians from 80 to 32, saving an estimated $380,000.

The key is showing the math: docs-to-GB factor, task rates, and scenario comparison. A vague ‘it’s too expensive’ fails; a defensible spreadsheet wins.

Deep Dive: Email Threading and Attachment Expansion Math

Email is the largest component in most matters. A single threaded conversation of 12 messages with 3 attachments can count as 1 document in business terms but 12+3×child items = 15 or more billable items. In a 2023 matter, we had 600,000 ’emails’ that expanded to 1.05 million items after processing. If your per-doc rate is $0.15, that difference is $67,500 versus $157,500.

Threading analysis reduces review volume by grouping, but processing still counts each message. Vendors may charge email threading as a separate task ($0.03/item). Model both: item count for storage, threaded count for review hours. I keep two columns in my template: ‘Items’ and ‘Review Units.’

How Data Reduction Technologies Change the Equation

Predictive coding and continuous active learning (CAL) can cut review volume 40–60%, but they add software fees or hourly training time. In a 2021 patent case, CAL reduced reviewed docs from 400k to 160k, saving $36,000 attorney review but costing $8,000 platform AI fee. Net positive, but only if your per-doc processing is already optimized.

Be cautious: data reduction does not shrink storage GB; processed data remains. So per-GB hosting unchanged. This is why a combined model often emerges: per-GB hosting + per-doc processing + hourly review with CAL.

Vendor Contract Fine Print That Affects Volume Cost

Beyond headline rates, these clauses matter:

  • Ingestion fee: flat $1,500–$5,000 regardless of volume; amortize across docs.
  • Minimum monthly storage: e.g., 100 GB minimum even if you use 20 GB.
  • Data export charge: $0.02–$0.05 per doc to produce final load files.
  • Re-processing fee: if you add custodians later, some vendors charge full rate again.

I once signed a per-doc deal with a $3,000 ingestion fee; on a small 40k doc case that equated to $0.075/doc hidden. Always normalize fixed fees into per-doc or per-GB equivalent in your model.

Sample Excel Formulas to Operationalize the Method

For those building the template, these exact formulas (assuming rows) help:

  • Weighted KB: =SUMPRODUCT(B2:B6,C2:C6) where B is mix, C is avg KB.
  • Total Raw GB: =WeightedKB*DocCount/1024/1024.
  • Processed GB: =RawGB*1.18 (overhead).
  • Per-GB Cost: =ProcessedGB*RateGB + ScannedPages*OCRrate.
  • Per-Doc Cost: =DocCount*RateDoc + ThreadingItems*ThreadRate.
  • Confidence Interval: =CONFIDENCE.NORM(0.05,StDev,SampleSize).

These mirror the logic in our Discovery Document Volume Cost Calculator but give you full transparency.

Common Mistakes That Blow Up Discovery Budgets

From experience, these recur and are absent from lightweight guides:

  • Counting ‘documents’ as narrative files but vendors bill ‘items’ including attachments and child objects.
  • Ignoring processing overhead (index +10–20%).
  • Assuming privilege review is only attorney time; platform charges per tag or per doc flagged.
  • Locking per-doc pricing when a late data dump doubles volume.
  • Forgetting production format (TIFF+LOADFILE costs more than native).
  • Using industry average docs/GB instead of your own sampled factor.

When I first built a budget for multistate litigation, I missed TIFF conversion for 20% of docs; that added $14,000 unexpected. Plan for format explicitly and revisit monthly.

When to Switch Pricing Models Mid-Case

Trade-offs are real. If your initial per-doc contract hits a volume cliff (e.g., 500k docs), renegotiate to per-GB for overflow. Most vendors allow mid-matter conversion with 30 days notice. But be wary: per-GB contracts often have minimum monthly commitments; if review slows, you pay for empty storage.

Honest limitation: no model eliminates risk. Hourly gives control but requires active management; fixed models need accurate forecasts. I have seen firms anchor to a model out of convenience and eat 30% overage. Build a trigger: if actual docs/GB deviates >15% from forecast, reopen pricing.

Final Checklist for Defensible Cost Estimation

Before you submit a discovery budget, verify:

  • Document-to-GB factor derived from actual custodian sample, not vendor average.
  • Hidden tasks (OCR, email threading, redaction) line-itemed with unit rates.
  • Three pricing scenarios computed side-by-side (per-doc, per-GB, hourly).
  • Pre-collection custodian sample performed with confidence interval noted.
  • Privilege review platform fees confirmed in writing.
  • Production format cost included (native vs image).
  • FRCP 26(b)(2)(B) proportionality argument prepared if scope challenged.

Following this DIY method answers how to calculate discovery document volume cost with numbers you can defend to a CFO or a magistrate. The work is front-loaded, but I have saved clients six figures by catching scope before collection. The spreadsheet you build today is the best protective order you will ever file.

Leave a Reply

Your email address will not be published. Required fields are marked *