· — Rishi Lahoti ·Oct 5, 2026·13 min read

Air-Gapped OCR: Document Extraction with No Internet Egress

What it actually takes to run air-gapped OCR in production — the weight supply chain, the offline update path, proving zero egress to an auditor, and why on-prem OCR no longer costs you accuracy.

Key Takeaways

  • Air-gapped OCR means every stage — layout detection, text recognition, and any model-based interpretation step — runs inside your network boundary with no outbound internet route, verifiable by cutting the network and watching the pipeline keep working
  • The accuracy penalty for staying private is effectively gone: PaddleOCR's PP-OCRv6_medium reaches 83.2% recognition accuracy and 86.2% detection Hmean with 34.5M parameters, beating Qwen3-VL-235B by 8.3 points on recognition at roughly 6,800× fewer parameters (PP-OCRv6 technical report)
  • The hard part is not inference — it's the model weight supply chain, the offline update path, and proving egress is actually zero to an auditor who won't accept "we configured it that way"
  • PaddleOCR ships under Apache 2.0, so code and weights can be mirrored into an isolated network and used commercially with no per-document licence fee and no vendor call-home
  • Prodinit built a fully air-gapped document extraction pipeline for a document AI client — PaddleOCR for text and layout, Qwen2.5-VL served via Ollama on NVIDIA NIM for structured interpretation, with zero network egress from the inference layer

Most deployments that call themselves "on-prem OCR" are not air-gapped. The models run locally, and then the stack quietly resolves a licence server, checks for a model update, or ships usage telemetry to a vendor endpoint — and the first person to notice is the auditor asking for six months of egress logs. That gap matters, because the regimes that force the question in the first place are about where data can travel, not where the GPU sits.

Air-gapped OCR runs document extraction entirely inside a network with no outbound internet route: model weights pre-loaded from an internal mirror, inference on private compute, and no vendor API call at any stage. In 2026 the accuracy cost is near zero — open-weight OCR models at 34.5M parameters now outscore frontier hosted vision-language models on OCR benchmarks.

What Air-Gapped OCR Actually Requires

Air-gapped OCR requires four things ordinary on-prem deployment does not: model weights resident on disk before the network is cut, an inference runtime with no licence check or telemetry callback, every supporting service reachable only over private networking, and a documented update path that moves new weights in through a controlled transfer rather than a download.

  1. Weights resident before isolation. The models — detection, recognition, layout, and any vision-language stage — are downloaded once in a staging environment, checksummed, and stored in an internal registry or object store. Nothing in the production path resolves a public hostname on first use. This is the single most common failure: toolkits that lazily fetch weights on first inference work perfectly in staging and break the moment they're deployed behind the air gap.
  2. A runtime with no call-home. Self-hosted does not mean offline. Audit the runtime for licence validation, update checks, crash reporting, and anonymous usage metrics, and disable each one explicitly rather than relying on the firewall to swallow it. A blocked callback still shows up as an outbound connection attempt in the egress log, and you will be asked to explain it.
  3. Private networking for every dependency. The document store, the extraction queue, the output database, the container registry, and the metrics backend all need private endpoints. On AWS this is the VPC interface endpoint inventory; the air-gapped LLM deployment guide covers the endpoint-gap diagnosis process in detail, and a missing endpoint is how most "zero egress" claims break.
  4. An offline update path. OCR models improve fast — PP-OCRv6 landed in June 2026, roughly a year after PP-OCRv5 — so "we'll never update" is not a plan. You need a repeatable transfer procedure with checksum verification and a rollback, documented before go-live.

The acceptance test is blunt and worth running in front of whoever signs off: disconnect the egress path entirely, run a representative batch of documents, and confirm the pipeline produces identical output. Anything that degrades or hangs was never air-gapped.

The Accuracy Trade-Off Disappeared in 2026

For years the honest pitch for on-prem OCR involved a trade: you keep your documents, you give up accuracy against hosted frontier models. That trade is no longer real. PaddleOCR's PP-OCRv6, released June 2026, ships tiers at 1.5M, 7.7M, and 34.5M parameters across 50 languages — and the mid tier outperforms hosted frontier systems on OCR benchmarks outright.

PP-OCRv6_medium reaches 83.2% recognition accuracy and 86.2% detection Hmean, improving on PP-OCRv5_server by 5.1 points on recognition and 4.6 points on detection. Against the hosted models, it beats Qwen3-VL-235B by 8.3 points on recognition while using roughly 6,800× fewer parameters, and exceeds Gemini-3.1-Pro on detection Hmean by 39.4 points (PP-OCRv6 technical report). The previous generation had already closed most of the gap: PP-OCRv5 added 13 percentage points over PP-OCRv4 on multi-scenario evaluation and cut recognition error on non-standard handwriting by 26% (PaddleOCR 3.0 Technical Report).

Two consequences follow for anyone building the business case. First, the accuracy objection to air-gapping a document pipeline is now a factual error rather than a trade-off, and should be challenged when a vendor raises it. Second, the hardware bill is small: a 34.5M-parameter detection-and-recognition stack is not a GPU-cluster workload, and PP-OCRv6 reports up to 5.2× faster CPU inference with OpenVINO, which means a meaningful share of air-gapped OCR deployments need no GPU at all. That changes the procurement conversation from "data centre capacity" to "a few servers you already own."

Where hosted frontier models still contribute is the semantic layer above character recognition — deciding which extracted number is the invoice total versus the subtotal, or whether a handwritten annotation overrides a printed value. That interpretation step is a separate architectural choice, covered in our intelligent document processing with VLMs breakdown, and it can be served privately too.

Getting Model Weights Into an Isolated Network

The weight supply chain is where air-gapped OCR projects actually slip, because it's a procurement and security-review problem disguised as an engineering task. Open weights must be fetched externally, scanned, approved, checksummed, mirrored to an internal registry, and version-pinned — and that chain needs an owner and an audit record, not a one-off download.

The licence position is the easy part. PaddleOCR is distributed under Apache 2.0, which permits commercial use, modification, and redistribution without royalty, so the weights can be mirrored internally and embedded in a commercial product with no per-document fee and no entitlement service to reach. That is a material difference from commercial OCR engines, where the air-gapped tier is often a separate SKU with an offline licence file that still has to be renewed through a manual process. Check the licence on each model separately — a toolkit under Apache 2.0 can ship pipelines whose individual checkpoints carry different terms.

A workable transfer pattern looks like this:

  1. Fetch and pin in a staging environment with internet access, recording the exact model version, file hashes, and upstream source.
  2. Scan and review the artefacts under the same process any third-party binary goes through — malware scanning plus a dependency and licence review.
  3. Mirror to an internal registry. Container images to a private registry; weights to internal object storage. The production environment pulls only from these.
  4. Pin by digest, not by tag. latest in an air-gapped environment is a silent-failure generator; an image digest is reproducible and auditable.
  5. Pre-bake weights into the image where image size allows, so no runtime fetch exists to fail.
  6. Document rollback — the previous pinned digest, and the procedure to revert if an updated model regresses on your document mix.

Hold back a labelled evaluation set inside the boundary and re-score it on every model update. Benchmark numbers are measured on public document mixes, and your invoices, claim forms, or scanned contracts are not that mix. An update that gains four points on OmniDocBench can still lose accuracy on your specific layouts, and without an internal eval set you'd have no way to know before it reached production.

Proving Zero Egress to an Auditor

"Zero egress" is a claim about evidence, not configuration. Auditors under CMMC and NIST SP 800-171 — whose Revision 3 defines 17 families of requirements for protecting controlled unclassified information — are not satisfied by a screenshot of a security group. They want controls making egress architecturally impossible, plus logs demonstrating none occurred.

Four artefacts carry most of that weight:

  • No route, not just no permission. Private subnets with no internet gateway and no NAT gateway means there is no path to block, rather than a path with a deny rule somebody can later amend. This is the difference between a control and a setting.
  • Flow logs with a documented review cadence. VPC flow logs (or the on-prem network equivalent) retained for the required period, with a named owner who reviews them and an alert on any outbound connection attempt. The alert firing on a runtime's blocked telemetry callback is exactly what you want to be able to show you caught.
  • A pinned, reproducible artefact inventory. Which model versions and image digests are running, where they came from, and who approved them — this is also what makes an extraction defensible months later when someone asks which model produced a specific output.
  • Per-document audit logging. Model version, timestamp, input document identifier, and output schema for every extraction, retained inside the boundary. For regulated document workflows this is frequently a hard requirement rather than good practice.

The architectural pattern is identical to any zero-egress AI deployment, which is why it's worth reusing rather than reinventing: Prodinit delivered a fully air-gapped AWS EKS platform for a regulated fintech in four weeks — no internet egress, 10+ VPC interface endpoints, private ECR — documented in the air-gapped EKS case study. An OCR pipeline drops into that topology as another private workload.

On-Prem OCR vs Cloud OCR: What You Actually Trade

With the accuracy gap closed, the on-prem OCR decision reduces to a short list of real trade-offs — operational ownership, elasticity, and who carries the compliance paperwork — rather than a quality compromise. The table below is the version worth putting in front of a security review: it separates what genuinely differs from what no longer does.

DimensionAir-gapped / on-prem OCRCloud OCR API
Document data boundaryNever leaves your networkTransits a vendor endpoint; governed by contract
OCR accuracy (2026)Competitive or better — PP-OCRv6_medium beats Qwen3-VL-235B on recognitionStrong, but no longer a clear lead
Cost shapeCapital plus operations; flat at volumePer-page or per-document; scales linearly forever
ElasticityBounded by hardware you provisionedEffectively unbounded
Operational burdenYours — serving, updates, monitoring, evalVendor's
Model update cadenceManual, controlled, auditableAutomatic, and sometimes without notice
Compliance paperworkArchitectural control; no transmission path to governBAA / DPA scope, subprocessor review, residency terms
Licence exposureApache 2.0 for PaddleOCR — no per-document feeMetered pricing, often with an air-gapped SKU premium

Air-gapped OCR is the right call when the documents themselves are the regulated asset — medical records, financial statements, contracts, CUI-bearing drawings — or when per-page pricing at your volume outruns the cost of running it yourself. Cloud OCR remains the better default for spiky, low-sensitivity, low-volume workloads where operational ownership is the dominant cost. For the broader buying decision across model serving as well as extraction, our private LLMs for regulated industries guide covers the same ground at platform level.

How Prodinit Builds Air-Gapped OCR Pipelines

Prodinit built a fully air-gapped document extraction pipeline for a document AI client: PaddleOCR handling text extraction and layout detection, with Qwen2.5-VL served via Ollama on NVIDIA NIM for structured interpretation, running inside a private on-premises environment with zero network egress from the inference layer. No document, prompt, or model response left the network at any point.

Three decisions from that build generalise. Ollama was the right serving choice because document processing throughput is measured per document rather than by simultaneous concurrent users, which sits well below the concurrency level where vLLM's continuous batching earns its added operational complexity — the full runtime comparison is in our on-prem LLM deployment guide. Pairing it with NVIDIA NIM's optimised inference path recovered throughput headroom without adopting a heavier stack. And splitting the work — cheap, debuggable OCR for extraction, a vision-language model only for semantic interpretation — kept both the compute bill and the failure surface smaller than a VLM-only pipeline reading raw pixels on every page.

The part worth planning for early is the eval set — build it before go-live, not after the first model update regresses something.

Get Prodinit's AI engineering guides in your inbox

Deep-dives on production LLMs, voice AI, and MLOps — published weekly. No sales emails.

Frequently Asked Questions

Air-gapped OCR is document text extraction where every stage — layout detection, character recognition, and any model-based interpretation — runs on private compute with no outbound internet route. Model weights are pre-loaded from an internal mirror, and no vendor API, licence server, or telemetry endpoint is contacted. The practical test is disconnecting the network and confirming output is unchanged.

Yes, and often better. PaddleOCR's PP-OCRv6_medium reaches 83.2% recognition accuracy and 86.2% detection Hmean with 34.5M parameters, beating Qwen3-VL-235B by 8.3 points on recognition and exceeding Gemini-3.1-Pro on detection Hmean by 39.4 points (PP-OCRv6 technical report). The accuracy argument against air-gapping a document pipeline no longer holds.

Through a controlled offline transfer, not a download. Fetch and pin the model in a staging environment, record file hashes, run security and licence review, mirror the artefacts to an internal registry, and pin production by image digest rather than tag. Re-score a labelled internal evaluation set before promoting any new version, and keep the previous digest for rollback.

Neither mandates an air gap. HIPAA permits processing PHI with cloud AI services under a Business Associate Agreement, and CMMC Level 2 aligns to the 110 controls in NIST SP 800-171 without specifying isolation. Air-gapping is chosen because it removes the transmission path a BAA exists to govern, converting a contractual control into an architectural one.

Yes. PaddleOCR is distributed under Apache 2.0, which permits commercial use, modification, and redistribution without royalty, and the models run entirely locally once weights are present on disk. Mirror the weights to internal storage and pre-bake them into your container image so nothing fetches at runtime, and verify the licence on each individual checkpoint you deploy.

Stay ahead in AI engineering.

Get the latest insights on building production AI systems, be the first to explore approaches that actually work beyond the demo.

Start a Project →