· — Dishant Sethi ·Oct 1, 2026·14 min read

AWS Bedrock vs OpenAI API: Choosing an LLM Provider

A practitioner's comparison of AWS Bedrock vs the OpenAI API — network isolation and VPC endpoints, compliance paperwork, cost structure, per-Region quotas, and model recency, from running Bedrock inside a zero-egress EKS platform.

Key Takeaways

  • AWS Bedrock vs OpenAI API is decided by where your data is allowed to travel and whose name goes on the contract — not by which model scores higher on a public benchmark. Both clear the capability bar for almost every production workload
  • Bedrock reaches inference over a VPC interface endpoint, so requests never touch the public internet; the OpenAI API is a public HTTPS endpoint, and no amount of network design changes that
  • Both providers now sell the same cost levers — roughly 50% off for asynchronous batch work and steep discounts on cached prompt prefixes — so the real difference is commitment shape: Bedrock's Provisioned Throughput reserves hourly Model Units, OpenAI's discounts stay per-token
  • Bedrock's operational tax is regional. On-demand throughput is capped by per-model, per-Region requests-per-minute and tokens-per-minute quotas, and model availability lags between Regions (AWS re:Post)
  • Prodinit ran Amazon Bedrock behind a VPC interface endpoint inside a fully air-gapped EKS platform for a regulated fintech, delivered in 4 weeks with zero internet egress (case study) — the provider decision was made by the network diagram, not the leaderboard

Most AWS Bedrock vs OpenAI API comparisons open with a model quality table. That table decides almost nothing. Both providers serve models that are more than capable enough for summarisation, extraction, classification, RAG answering, and agent loops — the workloads that make up the overwhelming majority of production systems.

The decisions that actually break a build come later: whether your prompts are allowed to leave your VPC, whether legal will sign the vendor's paperwork, and whether the throughput you tested in one AWS Region survives real traffic.

AWS Bedrock vs OpenAI API comes down to three axes: the data boundary, the procurement and compliance surface, and operational cost shape. Bedrock wins when inference must stay inside your own AWS network and on your existing AWS contract. The OpenAI API wins when you want the newest frontier models on the day they ship, with the shortest path from idea to running code.

Why AWS Bedrock vs OpenAI API Is Not a Model-Quality Decision

Choosing between AWS Bedrock and the OpenAI API on model quality assumes the two catalogues are disjoint. They largely are not. Bedrock serves Anthropic's Claude, Meta's Llama, Mistral, and Amazon's Nova families, plus OpenAI's open-weight gpt-oss-120b and gpt-oss-20b models with a 128K context window (AWS). For most tasks, several of these are interchangeable.

What is not interchangeable is the delivery mechanism. One of these providers hands you a model behind an IAM-authenticated endpoint that can be pulled entirely inside your own VPC. The other hands you an API key and a public hostname. Everything downstream — compliance evidence, network architecture, incident response, vendor risk review — follows from that single difference.

There is one genuine capability gap, and it runs in OpenAI's favour: OpenAI's proprietary frontier models and its newest interfaces land on OpenAI's own API first. Bedrock carries OpenAI's open-weight models, not the flagship closed ones. If your product depends on a specific OpenAI capability — the Realtime API for speech-to-speech, for instance — Bedrock is not an alternative at all, and the comparison ends there.

Feature-by-Feature Comparison

The table below covers the axes that change a production architecture, not the marketing feature matrix. Network isolation and quota management are the two rows that have redirected the most Prodinit engagements; raw model quality has never been the deciding row.

DimensionAmazon BedrockOpenAI API
Model catalogueClaude, Llama, Mistral, Amazon Nova, OpenAI open-weight gpt-ossOpenAI frontier and small models only
Model recencyThird-party models arrive after their native launchNewest OpenAI models on day one
Network pathVPC interface endpoint (PrivateLink) — no internet egress requiredPublic HTTPS endpoint
AuthenticationAWS IAM / SigV4, IRSA for pods on EKSBearer API key
Data retention defaultNo customer prompt storage for inference; governed by your AWS agreementUp to 30 days by default; zero data retention (ZDR) on eligible endpoints by request (OpenAI)
Compliance postureHIPAA-eligible under the AWS BAA, in-scope for SOC, PCI DSS and FedRAMP programmesSOC 2 Type 2; BAA available, scoped to ZDR-eligible endpoints
Throughput controlPer-model, per-Region RPM and TPM quotas; Provisioned Throughput in hourly Model UnitsAccount- and tier-based rate limits that rise with usage history
Batch economics~50% off on-demand, results to S3 within 24h~50% off via Batch API (24h) or the Flex service tier
Regional behaviourModel availability varies by Region; cross-Region inference profiles spread loadGlobally routed, no Region to select
BillingConsolidated on the AWS bill, counts toward existing AWS commitmentsSeparate vendor, separate invoice, separate procurement
Guardrails / retrievalBedrock Guardrails and Knowledge Bases built inBuild it yourself or bring a framework
Best-fit problemRegulated data, AWS-native stacks, consolidated procurementFastest access to frontier capability

The Data Boundary Decides More Than the Benchmark

The data boundary is the first question to answer, because it eliminates one provider outright in regulated builds. Amazon Bedrock is reachable through a VPC interface endpoint, so an EKS pod on a private subnet with no internet gateway can still call a foundation model. The OpenAI API has no equivalent — reaching it requires a route out of your network.

Prodinit built a fully air-gapped AWS EKS platform for a regulated fintech in four weeks: 10+ VPC interface endpoints, private ECR, no NAT gateway on private subnets, and Amazon Bedrock as the inference path. Zero internet egress was a hard requirement from the first architecture review. On that build, no model benchmark was ever consulted — the shortlist was one provider long before model selection began. The air-gapped LLM deployment guide covers the VPC endpoint configuration in detail, and the air-gapped AI for fintech post covers the five infrastructure layers underneath it.

If your constraint is HIPAA rather than full network isolation, both providers are viable, but the paperwork differs in shape. Bedrock inherits the AWS BAA you likely already have. OpenAI signs a BAA on request, but it is scoped to endpoints eligible for zero data retention — several convenience features fall outside that scope, which means an engineer can take PHI outside the agreement simply by calling the wrong endpoint. That is an architecture review item, not a contract item. Our HIPAA-compliant LLM deployment post covers the controls that make the distinction enforceable.

Cost Structure: Compare Commitment Shapes, Not Sticker Prices

Per-token list prices move constantly and both providers discount aggressively, so comparing published rates on the day you evaluate tells you very little about your annual bill. Compare the shape of the commitment instead. Bedrock and OpenAI both offer roughly 50% off for asynchronous batch processing and large discounts on cached prompt prefixes; what differs is what happens when you want guaranteed capacity.

Bedrock's reserved path is capacity, not tokens. Provisioned Throughput buys Model Units billed per hour on a one-month or six-month commitment, charged whether or not traffic arrives. It is also mandatory to serve a model you fine-tune on Bedrock. That converts a variable cost into a fixed one — good for steady, high-volume traffic with a latency SLA, punishing for spiky workloads.

OpenAI's reserved path stays per-token. The Batch API returns results within 24 hours at roughly half the on-demand rate, and the Flex service tier offers the same discount on synchronous-style calls in exchange for variable latency. Both stack with prompt caching, which discounts repeated prefixes heavily (OpenAI) — the biggest single saving available to agent loops and long system prompts, as the prompt caching explainer sets out.

Two second-order costs get missed in nearly every evaluation. VPC interface endpoints bill per hour per availability zone plus per GB processed, so the Bedrock network layer has a floor cost before a single token moves. And on the OpenAI side, a separate vendor invoice means separate procurement, separate security review, and spend that does not count toward an existing AWS commitment — a real number for enterprises on a negotiated AWS agreement. Our LLM cost optimisation post covers how to model the total rather than the rate.

The Operational Differences That Surface After Launch

Three operational differences consistently surprise teams in the first month of production, and all three sit on the Bedrock side: per-Region throughput quotas, uneven model availability between Regions, and cross-Region routing. All three are manageable, but only if you design for them before launch rather than diagnosing them during an incident.

Per-Region quotas throttle before you expect. Bedrock enforces separate requests-per-minute and tokens-per-minute limits per model, per Region, concurrently. A workload that tested cleanly returns ThrottlingException: Too many requests the first time real traffic lands (AWS re:Post). Raise quotas through Service Quotas before launch, not after.

Model availability is regional. A model generally available in us-east-1 may not be enabled in eu-central-1 for weeks, and some never arrive. If data residency pins you to a Region, verify the specific model is served there before you design around it.

Cross-Region inference exists because one Region often is not enough. Bedrock inference profiles route a single invocation across multiple Regions to raise effective throughput (AWS docs). That is a useful capacity tool and a compliance decision at the same time: routing across Regions moves data across them, which can undo the residency guarantee that pushed you to Bedrock in the first place.

The OpenAI API trades these for a different failure mode. There is no Region to select and no quota console to pre-warm — rate limits rise with account history, which is simpler until a launch spike arrives faster than your tier does. Either way, the instrumentation requirement is identical: log model, latency, token counts and cost per call from day one, as the LLM observability post sets out.

When Amazon Bedrock Is the Right Choice

Choose Amazon Bedrock when the network boundary, the compliance surface, or the procurement path is the binding constraint rather than model capability. In practice that covers most regulated and enterprise builds — healthcare, finance, government and defence — and the decision is usually settled before any model evaluation begins.

  • Prompts or outputs cannot leave your network. The VPC interface endpoint is the whole argument. The only alternative is self-hosting open-weight models yourself, which costs considerably more to operate.
  • You are already deep in AWS. IAM roles, IRSA on EKS, CloudWatch, Secrets Manager and KMS all work without a second identity system or a key to rotate by hand.
  • Procurement is the slow path. Inference on the existing AWS bill avoids a new vendor review and counts toward a negotiated commitment.
  • You want model optionality. Switching between Claude, Llama, Mistral and Nova is a model ID change, not a new integration — which matters when a cheaper model becomes good enough for a sub-task.
  • You need built-in guardrails and retrieval. Bedrock Guardrails and Knowledge Bases cover ground you would otherwise assemble yourself.

Prodinit's AI infrastructure and LLMOps practice builds this layer regularly, and the pattern generalises past fintech — the same architecture underpins the self-hosted LLM playbook for enterprise.

When the OpenAI API Is the Right Choice

Choose the OpenAI API when capability recency and iteration speed matter more than network isolation, which describes most pre-product-market-fit builds and almost every internal tool. The integration is a key and an SDK, and the newest models are available the day they launch.

  • You need a specific OpenAI capability. The Realtime API, the newest reasoning models, and the latest interfaces are not on Bedrock. If your product depends on one, there is no comparison to run.
  • You are still finding the product. Shipping a prototype in an afternoon beats a correct long-term architecture you have not validated demand for.
  • Your data is not regulated. Public documentation, internal knowledge, marketing content and generated code rarely justify the isolation overhead.
  • Your traffic is spiky. Per-token billing with no hourly commitment is strictly better than reserved Model Units sitting idle overnight.

The honest default for an early-stage team is the OpenAI API with a provider abstraction in front of it — and a plan for the migration that the first enterprise security questionnaire will trigger.

Build the Abstraction Before You Pick

Whichever provider you choose, put it behind an interface. The AWS Bedrock vs OpenAI API decision has a shelf life measured in quarters: pricing moves, a model you depend on is deprecated, a new customer arrives with a data residency clause, and the right answer changes without your architecture getting a vote.

Keep the provider-specific code in one adapter that owns authentication, retries with exponential backoff, streaming, token accounting and structured-output parsing. Everything above it speaks a single internal interface. Prodinit built Cuebo's voice pipeline with five swappable AI pipeline variants behind a factory for exactly this reason (case study) — the provider became a configuration value rather than a rewrite.

Two details make that abstraction real rather than theoretical. Normalise errors early, because a Bedrock ThrottlingException and an OpenAI 429 need the same backoff behaviour from your caller. And run your own evaluation set against both providers before committing — public benchmarks are run on datasets that look nothing like your prompts, and the gap between two capable models on your traffic is usually smaller than the gap between two prompt revisions.

Get Prodinit's AI engineering guides in your inbox

Deep-dives on production LLMs, voice AI, and MLOps — published weekly. No sales emails.

Frequently Asked Questions

Not reliably, and the comparison is unstable because both providers change list prices often. Per-token rates land in the same range for comparable models. The structural difference is commitment: Bedrock's Provisioned Throughput bills reserved Model Units hourly whether or not traffic arrives, while OpenAI stays per-token with Batch and Flex discounts. Model your actual traffic shape against both pricing pages.

No. The OpenAI API is a public HTTPS endpoint, so a private subnet with no route out cannot reach it. The options are a controlled egress path through a NAT gateway and allowlisting proxy, which reintroduces the egress you were avoiding, or Amazon Bedrock over a VPC interface endpoint. For zero-egress requirements, Bedrock or self-hosted models are the only paths.

Partly. Bedrock serves OpenAI's open-weight models — gpt-oss-120b and gpt-oss-20b, both with a 128K context window — available to Bedrock users without a separate access request. OpenAI's proprietary frontier models and newer interfaces such as the Realtime API remain exclusive to OpenAI's own platform and, for some models, Azure OpenAI.

Both can be made compliant, but Bedrock has less room for error. Bedrock is HIPAA-eligible under the AWS BAA most healthcare organisations already hold. OpenAI signs a BAA on request, scoped to endpoints eligible for zero data retention, so an engineer can take PHI outside the agreement by calling a non-eligible endpoint. Enforce the boundary in architecture, not policy.

Answer three questions in order. Can prompts and outputs leave your network? If no, Bedrock. Does the product depend on a specific OpenAI-only capability such as the Realtime API? If yes, OpenAI. If neither binds, choose on operational fit — existing AWS footprint and procurement favour Bedrock, iteration speed and model recency favour OpenAI — and put both behind one adapter.

Stay ahead in AI engineering.

Get the latest insights on building production AI systems, be the first to explore approaches that actually work beyond the demo.

Start a Project →