All posts
16 min read

Lambda Concurrency: Reserved vs. Provisioned — The DVA-C02's Most Counterintuitive Trade-Off

Reserved concurrency and provisioned concurrency sound like upgrades of the same feature. They are not — and the DVA-C02 exploits exactly that confusion. Here's the complete guide to getting every Lambda concurrency question right on exam day.

Here is a scenario you will see on the DVA-C02:

A developer notices that a critical payment processing Lambda function occasionally experiences delays of 500–800 ms on the first invocation after a period of inactivity. The function is written in Java and handles latency-sensitive transactions. Which action should the developer take to eliminate these delays?

Four options appear. One says increase the Lambda function's memory allocation. One says configure reserved concurrency of 10. One says enable provisioned concurrency of 5. One says set a higher function timeout.

The answer is provisioned concurrency — but if you chose reserved concurrency because it "sounds more permanent" or "like it keeps instances warm," you're not alone. This is one of the highest-frequency conceptual traps on the DVA-C02. And the confusion runs deep: reserved concurrency and provisioned concurrency both have "concurrency" in the name, they're configured in the same part of the AWS Lambda console, and they sound like two tiers of the same feature. They are fundamentally different mechanisms that solve completely different problems.

This guide gives you the deep architectural understanding — not just "provisioned = warm" — that lets you decode every variant of this question in under 90 seconds.


Why Lambda Cold Starts Happen (and Why Memory Doesn't Fix Them)

Before comparing the two concurrency types, you need to understand what a cold start actually is — because the DVA-C02 tests this at the architectural level, not just the symptom level.

When a Lambda function is invoked, one of two things happens:

  1. Warm invocation: An existing execution environment — a micro-VM with your code already loaded — picks up the request. Response starts in milliseconds.
  2. Cold start: No existing environment is available. Lambda must provision a new execution environment, download your deployment package, start the language runtime (the JVM for Java, the Node.js process for Node, etc.), and run your initialization code before your handler ever runs.

For a small Node.js function with minimal dependencies, a cold start might add 100–200 ms. For a Java function using Spring Boot or a large dependency tree, cold starts routinely run 1–5 seconds. This is not a bug — it is the architecture's price for infinite scale. Lambda does not keep a pool of pre-warmed execution environments running for every function forever. That would defeat the entire cost model of serverless.

Increasing memory allocation makes execution faster once the runtime is running — it does not affect environment provisioning time. A 3008 MB Java function still does a full JVM startup on cold start; it just computes faster afterward. Memory is irrelevant to cold start latency. Timeout is also irrelevant — it only controls how long a function is allowed to run, not startup behavior.

This is the first trap the exam sets: it presents memory as the obvious "performance lever" and bets that you'll reach for it before understanding what's actually slow.


Reserved Concurrency: A Cap, Not a Warmth Guarantee

Reserved concurrency is a limit, not a warmth feature. When you set reserved concurrency on a Lambda function, you are doing two things simultaneously:

1. Guaranteeing capacity. You carve out that number of concurrent executions from your account's regional concurrency pool (default: 1,000 per region) and dedicate them exclusively to that function. No other function can use those slots, even if demand spikes elsewhere. This prevents your payment function from being throttled because your image-processing function suddenly went viral.

2. Capping throughput. If your function receives more simultaneous invocations than its reserved concurrency, the excess requests are throttled. They receive a 429 TooManyRequests error (or, for synchronous invocations, the caller gets a throttle exception). Reserved concurrency does not queue the overflow — it rejects it.

What reserved concurrency does NOT do: it does not keep execution environments warm. If your function hasn't been invoked in 15 minutes and then receives a request, Lambda still has to cold-start a new environment. The fact that you reserved 10 slots means 10 cold starts can happen in parallel — not zero cold starts.

The real-world use case for reserved concurrency:

  • Protect a downstream dependency (a database, a rate-limited API) from burst traffic by capping Lambda throughput at its safe maximum.
  • Ensure a critical function always has capacity, immune to noisy-neighbor consumption from other functions in the account.
  • Implement traffic shaping for fan-out architectures.

None of these use cases involve eliminating cold starts. Reserved concurrency is a ceiling and a reservation, not a warmth mechanism.


Provisioned Concurrency: Pre-Initialized Environments

Provisioned concurrency solves the cold start problem directly. When you configure provisioned concurrency on a Lambda function version or alias, AWS pre-initializes that many execution environments and keeps them running continuously, even with zero traffic.

These environments run your initialization code (everything outside your handler function) before any request arrives. When an invocation comes in, it hits an already-warm environment and your handler executes immediately. From the caller's perspective, there is no startup penalty — the first request looks identical to the thousandth.

Provisioned concurrency costs money in a way that reserved concurrency does not. You pay per GB-second for the provisioned concurrency hours, even if the environments receive no traffic. You are essentially renting idle compute to guarantee response time. For a function that receives traffic around the clock, the cost is usually justified. For a batch job that runs at 2 AM, it probably is not.

The real-world use case for provisioned concurrency:

  • Latency-sensitive APIs where cold start variance is unacceptable (payment processing, real-time recommendations, voice response systems).
  • Functions using large runtimes (JVM, .NET) where initialization is inherently slow.
  • APIs that need to meet SLA commitments for P99 latency.

Provisioned concurrency can be combined with Application Auto Scaling, which adjusts the number of provisioned environments based on a schedule or on Lambda's reported metric (ProvisionedConcurrencyUtilization). This lets you pre-warm for known traffic spikes (a sale event, a morning rush) without paying for maximum provisioned capacity 24/7.


Side-by-Side Comparison

DimensionReserved ConcurrencyProvisioned Concurrency
Primary purposeCapacity reservation and throttle capEliminate cold starts
Eliminates cold starts?NoYes
Prevents throttling?Partially (reserves slots; excess still throttles)Only if set high enough
Additional cost?No ($0 beyond standard Lambda pricing)Yes (per GB-second, plus invocation cost)
Configures onFunction (any version or $LATEST)Function version or alias (not $LATEST)
Minimum value0 (disables the function entirely if set to 0)1
Effects when limit hitRequests are throttled (429)Overflow to on-demand concurrency (cold starts return)
Use with Auto Scaling?NoYes

Two entries in that table are exam traps on their own:

  • Reserved concurrency of 0 disables the function entirely. The exam will offer "set reserved concurrency to 0" as a way to "pause" a function during maintenance windows — and it is correct. This is surprising behavior worth memorizing.
  • Provisioned concurrency cannot be set on $LATEST. It requires a published version or an alias pointing to a version. This trips up candidates who work primarily from the console and never publish versions.

The Account-Level Concurrency Pool

Both mechanisms interact with your account's regional concurrency pool, and the DVA-C02 tests this at the systems level.

AWS accounts get 1,000 concurrent Lambda executions per region by default (this can be increased via a service quota request). Concurrency in Lambda means simultaneous in-flight requests — not invocations per second.

When you set reserved concurrency of 100 on Function A:

  • 100 slots are permanently set aside for Function A.
  • The remaining 900 are available to all other functions (unreserved pool).
  • If Function A is receiving only 20 concurrent requests, the other 80 reserved slots sit idle — they do NOT flow back into the unreserved pool. This is the cost of reservation.

When you set provisioned concurrency of 20 on Function A:

  • Those 20 environments count against your reserved concurrency (if configured) or the unreserved pool.
  • If requests exceed 20, Lambda spins up additional on-demand environments — subject to your reserved limit or the account pool.
  • You are charged for the 20 provisioned environments continuously.

The exam will ask you to calculate or reason about what happens when multiple functions with reserved concurrency are all at maximum simultaneously, or when a spike pushes total concurrency above the account limit. The mental model: reserved concurrency creates hard walls between function pools; the unreserved pool is shared and first-come-first-served.


Exam Scenario Walkthroughs

Scenario 1: The Payment API Cold Start Problem

A financial services company runs a payment validation API on Lambda (Java 17 runtime). P99 latency must be under 200 ms per their SLA. The development team reports occasional latency spikes of 2–3 seconds on the first request after an idle period. They currently have reserved concurrency set to 50. What should they add or change?

Option A: Increase reserved concurrency to 100.
Option B: Enable provisioned concurrency of 10 on the deployed alias.
Option C: Increase the Lambda function memory to 3008 MB.
Option D: Enable function URL with streaming response.

Answer: B. The team already has reserved concurrency — they are not seeing throttling, they are seeing cold starts. Increasing reserved concurrency (Option A) just carves out more idle slots; it changes nothing about warm-up time. More memory (Option C) makes execution faster post-startup, but the JVM startup itself is the bottleneck. Streaming response (Option D) is unrelated.

Provisioned concurrency of 10 pre-initializes 10 environments, eliminating cold start latency for up to 10 concurrent requests. If they occasionally see burst traffic above 10, those additional invocations would cold-start, but the baseline SLA-affecting scenario is solved.

Distractor anatomy: Option A feels right because "more reserved = more permanent capacity." The exam exploits the natural reading of "reserved" as meaning "pre-warmed." Reject it by returning to the core definition: reserved concurrency = capacity guarantee + cap, not warmth.


Scenario 2: The Noisy Neighbor Problem

A developer works at a company with multiple Lambda functions sharing an AWS account. A mission-critical order processing function is being throttled when an unrelated data analytics function spikes to 800 concurrent executions, consuming most of the account's concurrency pool. What is the most appropriate fix?

Option A: Enable provisioned concurrency on the order processing function.
Option B: Set reserved concurrency on the order processing function.
Option C: Move the analytics function to a separate AWS account.
Option D: Set a higher timeout on the order processing function.

Answer: B. This is the canonical use case for reserved concurrency. By reserving, say, 200 concurrent executions for the order processing function, those slots are guaranteed regardless of what the analytics function does. The analytics function is capped to the remaining unreserved pool.

Distractor anatomy: Option A (provisioned concurrency) is tempting because it sounds like a "stronger" solution. But provisioned concurrency does not protect against throttling in the way reserved concurrency does — it just pre-warms environments. If the account pool is exhausted and the order function has no reserved concurrency, provisioned concurrency's pre-warmed environments might still be available, but new invocations beyond those provisioned slots would throttle. Reserved concurrency is the correct lever for isolation. Option C (separate account) is a valid architectural pattern but is not the most appropriate fix — it's expensive and complex for this scenario.


Scenario 3: The Disabled Function Trap

A developer wants to completely stop all invocations of a Lambda function for an emergency maintenance window without deleting the function or its configuration. What is the fastest way to accomplish this?

Option A: Set the function's timeout to 1 second.
Option B: Remove all triggers from the function.
Option C: Set reserved concurrency to 0.
Option D: Delete the function's IAM execution role.

Answer: C. Setting reserved concurrency to 0 immediately throttles all invocations — the function exists, its code and configuration are intact, but no request can execute. Setting it back above 0 restores it instantly. This is the correct "pause" mechanism.

Distractor anatomy: Option B (removing triggers) only prevents event-source invocations — direct invocations via the SDK, CLI, or other services would still work. Option D is destructive and hard to reverse. Option A doesn't prevent invocations — it just causes them to time out after 1 second, which may still execute business logic and incur costs.


The Mental Model: Lock vs. Key

Here is the mental model that makes this stick under exam pressure:

Reserved concurrency is a lock on the door. It defines how many people can be inside the room at once and ensures nobody from the hallway (other functions) can take those spots. It does not heat the room in advance. The first person to walk in after a long absence still finds it cold.

Provisioned concurrency is a heater left on. It keeps the room at temperature so the first person who walks in finds it comfortable immediately. It does not lock out anyone — if more people arrive than the heater can handle, they enter a cold room (on-demand cold start).

You often want both: lock the room so the SLA-critical function always has space (reserved), and pre-heat it so the first request is fast (provisioned). These are not mutually exclusive. Provisioned concurrency counts against your reserved concurrency allocation if both are configured on the same function.

When you see "throttling problem" → reach for reserved concurrency.
When you see "cold start / latency problem" → reach for provisioned concurrency.

These two phrases are your exam diagnostic keys.


Common Traps to Avoid

Trap 1: Treating reserved concurrency as a warm-up feature.
The name "reserved" implies permanence, which the brain maps to "always available and warm." It is permanent capacity, but cold starts happen in that capacity.

Trap 2: Forgetting that reserved concurrency of 0 disables the function.
Zero is not a neutral no-op — it is an off switch. The exam uses this as a correct answer for "stop all invocations immediately."

Trap 3: Confusing provisioned concurrency with memory.
Both affect performance. Memory affects execution speed. Provisioned concurrency affects startup speed. For runtime-heavy functions (Java, .NET), startup dominates; for lightweight runtimes (Node.js, Python) with small packages, startup is negligible.

Trap 4: Applying provisioned concurrency to $LATEST.
You cannot configure provisioned concurrency on the $LATEST version. You must publish a version (aws lambda publish-version) and then configure provisioned concurrency on that version number or on an alias pointing to it. The exam may present this as a gotcha — if a developer tried to enable provisioned concurrency and received an error, the likely cause is that they targeted $LATEST.

Trap 5: Assuming provisioned concurrency prevents all throttling.
If your provisioned concurrency is set to 10 and you receive 50 simultaneous requests, 10 use provisioned environments and 40 spin up on-demand — subject to your reserved concurrency cap or the account pool. Provisioned concurrency doesn't grant unlimited capacity.

Trap 6: Not accounting for the unreserved pool drain.
If you set reserved concurrency across multiple functions that together exceed 1,000 (the default limit), the last function you configured may receive an error or others may be throttled unexpectedly. The quota is shared, and reserved concurrency permanently removes slots from the unreserved pool.


Exam-Day Checklist

Before leaving any Lambda concurrency question, verify:

  • Is the problem slow first request (cold start)? → Provisioned concurrency.
  • Is the problem throttling / 429 errors? → Reserved concurrency or account limit increase.
  • Is the scenario asking to isolate a function from noisy neighbors? → Reserved concurrency.
  • Is the scenario asking to disable a function temporarily? → Reserved concurrency = 0.
  • Does the scenario mention Java, .NET, or large dependency packages? → Cold start is more likely the culprit; lean toward provisioned concurrency.
  • Is the answer asking you to configure provisioned concurrency on $LATEST? → That's a trap — it requires a published version or alias.
  • Is the scenario asking to reduce cost while handling bursty traffic? → On-demand (no provisioned concurrency) is cheaper for irregular workloads; provisioned only makes economic sense for consistent baseline traffic.
  • Is the scenario about preventing a downstream service from being overwhelmed? → Reserved concurrency as a cap, not provisioned concurrency.

Putting It Together: When You'd Use Both

A well-architected production Lambda function handling latency-sensitive traffic might look like this:

  1. Publish a version of the function after each deployment.
  2. Configure an alias (e.g., prod) pointing to the latest version.
  3. Set reserved concurrency of 200 on the function — guaranteeing capacity in the account pool.
  4. Set provisioned concurrency of 20 on the prod alias — pre-warming 20 environments so P99 latency is predictable.
  5. Configure Application Auto Scaling on the provisioned concurrency, scheduling it to increase before known traffic peaks (e.g., 9 AM weekdays) and scale back down after.

In this setup, the first 20 concurrent requests always hit warm environments. Requests 21–200 experience cold starts but are protected from account-level throttling. Requests beyond 200 are throttled (and the upstream service should handle that with exponential backoff and retry). The combination solves both the isolation problem (reserved) and the latency problem (provisioned).

The DVA-C02 is a developer's exam — it expects you to know not just which feature to use but why, and how they compose. Memorizing "provisioned = warm" passes maybe half the questions. Understanding the architecture — concurrency pool, version requirements, cost trade-offs, interaction between the two features — is what gets you through the scenario variants.


Ready to Test Yourself?

If you got through this and want to see how DVA-C02 questions actually phrase this in the exam context, try the free DVA-C02 diagnostic on CertCoach. Ten questions across the four domains — no signup, instant per-question AI explanations showing you exactly why each wrong answer is wrong.

If you're preparing seriously, the Sprint Pass at $39 (code LAUNCH50 for 50% off) covers all three associate certs — DVA-C02, SAA-C03, and AIF-C01 — with adaptive mocks, the full question bank, and unlimited AI tutor access. One pass, all three exams.

For SAA-C03 readers who landed here: the SQS vs. SNS vs. EventBridge guide covers the messaging layer that often appears alongside Lambda questions in architect-level scenarios, and the Aurora vs. RDS deep dive is the database equivalent of this post — same depth, same exam-trap structure.

Lambda concurrency is one of those topics that feels solved once you understand it — and once you do, you'll see the same underlying question coming from five different angles on the actual exam. Get it right once, get it right every time.