All posts
13 min read

AWS Step Functions vs EventBridge: The Orchestration vs Choreography Decision the SAA-C03 and DVA-C02 Test

Step Functions and EventBridge both coordinate AWS services — but they solve opposite problems. Master the orchestration vs choreography distinction and decode every workflow question on the SAA-C03 and DVA-C02.

Here is a scenario you will see on the SAA-C03:

A media company needs to process uploaded videos through four sequential steps: transcode the video, generate thumbnails, update a database record, and send a notification to the publisher. If the transcoding step fails, the workflow should automatically retry up to three times with exponential backoff. A solutions architect must choose the most operationally efficient design. What should the architect recommend?

Four answer choices: Amazon SQS with four Lambda consumers chained together. Amazon EventBridge with four rules routing events between steps. AWS Step Functions Standard Workflow with Lambda states for each step. AWS Step Functions Express Workflow with Lambda states for each step.

The answer is Step Functions Standard Workflow — but only if you understand why each of the others fails, and how to distinguish Standard from Express. Many candidates who understand Step Functions still lose this question because they're unsure which workflow type fits the requirement for retry history and audit logging.

The aws step functions vs eventbridge decision is one of the most important architectural choices the SAA-C03 and DVA-C02 test — and the confusion runs deep because both services "coordinate AWS services" and both live under the broad umbrella of event-driven and serverless architecture. This guide gives you the precise mental model, the decision tree, and the exact signal words that make every variant of this question straightforward.


The Core Distinction: Orchestration vs Choreography

The concept that unlocks every Step Functions vs EventBridge question is orchestration vs choreography. These are two opposite patterns for coordinating multiple services in a distributed system.

AWS Step Functions orchestration vs EventBridge choreography architecture diagram for SAA-C03 and DVA-C02 — showing centralized workflow control vs event bus fan-out

Orchestration (Step Functions)

In orchestration, a central coordinator knows the entire workflow. It tells each service what to do, waits for the result, handles failures, and decides what happens next.

AWS Step Functions is the orchestrator. When you define a state machine, you are explicitly encoding:

  • What runs (Lambda functions, DynamoDB actions, ECS tasks, SageMaker jobs, API calls — over 220 AWS SDK integrations)
  • In what order (sequential or parallel states)
  • What happens on failure (Retry with configurable backoff; Catch blocks that route to compensating states)
  • What data flows between steps (InputPath, OutputPath, Parameters, ResultSelector — JSON transformations between states)

The coordinator is visible. Execution history is stored. You can open the AWS console, look at any execution, and see exactly which state it is in, how long it has been there, and what data it received and produced at every step.

Choreography (EventBridge)

In choreography, there is no central coordinator. Services emit events; other services listen for those events and react. The services do not know about each other — they know only about the event bus.

Amazon EventBridge is the choreographer. Services publish events to an event bus; rules match incoming events by JSON pattern and route them to target services. If three targets are registered for the same rule, all three receive the event simultaneously. No service tracks what happened to the overall business transaction — each just handles its own slice.

The bus is invisible. There is no built-in "workflow state" or execution history across the chain of reactions. Choreography is powerful for loose coupling and fan-out, but it can make debugging complex chains of events difficult.


AWS Step Functions: What the Exam Actually Tests

Standard Workflows vs Express Workflows

Step Functions offers two fundamentally different workflow types. Getting this wrong under exam pressure is easy — the names don't clearly signal the architectural difference.

DimensionStandard WorkflowExpress Workflow
Max durationUp to 1 yearUp to 5 minutes
Execution semanticsExactly-onceAt-least-once (async) or At-most-once (sync)
Execution historyFull, queryable, stored for 90 daysCloudWatch Logs only
ThroughputUp to 2,000 executions/second per accountUp to 100,000 executions/second
Pricing modelPer state transition ($0.025/1,000)Per execution duration + request count
Best forLong-running, auditable, exactly-once workflowsHigh-volume, short-lived, cost-sensitive tasks
Exam signals"audit log," "compliance," "run > 5 min," "human approval""IoT," "high-throughput," "streaming," "milliseconds"

The single most tested exam trap in this dimension: a candidate reads "retry with backoff" and thinks Express Workflow, because retries sound like a short-running reliability feature. But if the scenario also mentions "audit trail," "compliance requirement," or a workflow that can run longer than 5 minutes, the answer shifts to Standard Workflow. Standard Workflow's execution history IS the audit trail.

State Types (DVA-C02 Goes Deep Here)

The DVA-C02 tests Step Functions state types in detail. The SAA-C03 tests them at a conceptual level. Know these:

  • Task state — does work: invokes a Lambda, calls an AWS SDK action, or runs an ECS task via an optimized service integration. The most common state on the exam.
  • Choice state — branches based on a condition. The exam's equivalent of an if/else: "if the order total exceeds $1,000, route to the fraud-review state, otherwise proceed to payment."
  • Wait state — pauses for a fixed time or until a timestamp. Used when a downstream service needs time to process before the next step.
  • Parallel state — runs multiple branches simultaneously, then waits for all to complete before continuing. The exam signal: "these steps can run concurrently."
  • Map state — iterates over an array, running the same sub-workflow for each item. The exam signal: "process each item in a list independently."
  • Pass state — passes input to output (optionally transforming it). Used to inject test data or constant values into a workflow; rarely the exam answer.
  • Succeed / Fail states — terminal states that stop the execution successfully or with a named error.

Callback Pattern with Task Tokens (DVA-C02 Specific)

The callback pattern solves a specific problem: you need to pause a Step Functions execution and wait for an external system to complete before continuing — for example, waiting for a human to approve a transaction in a ticketing system, or waiting for a third-party payment API to call back when settlement completes.

The pattern: a Task state sends a task token (an opaque string generated by Step Functions) to the external service. The Step Functions execution pauses — it does not poll, does not consume resources, and does not timeout as long as its heartbeat setting allows. When the external service finishes, it calls SendTaskSuccess or SendTaskFailure with the token. Step Functions resumes.

The exam signal: "the workflow must pause and wait for an external actor or third-party service to respond before continuing." The specific answer choice will include "task token" and "SendTaskSuccess" somewhere in the description.


Amazon EventBridge: What the Exam Actually Tests

The Event Bus Model

EventBridge operates on a pub/sub model built around structured JSON events. Every event has a source, a detail-type, and a detail body. Rules match events by pattern — they can match on any combination of these fields, including nested values in the detail.

Example event from S3:

{
  "source": "aws.s3",
  "detail-type": "Object Created",
  "detail": {
    "bucket": { "name": "my-uploads" },
    "object": { "key": "videos/new-upload.mp4" }
  }
}

A rule matching source = aws.s3, detail-type = Object Created, and detail.object.key matching videos/* would route this event to a Lambda function that starts video processing — without any polling, without any infrastructure change, and without the Lambda needing to know what produced the event.

This is the EventBridge exam model: an AWS service emits an event; you write a rule to react. The key insight is that you never write code to detect AWS service state changes. EventBridge already receives them.

EventBridge Event Sources

Know these event source categories for the exam:

  • AWS services (default bus): Over 100 AWS services emit events to the default event bus automatically. S3, EC2, RDS, CodePipeline, AWS Config, IAM — all emit structured events you can match against without any configuration on the source side.
  • Custom applications (custom buses): Your application calls PutEvents to send custom events to a custom event bus. This is how you decouple application components — the order service emits OrderPlaced; the inventory service, email service, and analytics service each have rules that react.
  • SaaS partners (partner event sources): Zendesk, Datadog, PagerDuty, Stripe, and dozens more send events directly to EventBridge without any webhook configuration on your side. The exam signal: "integrate with a SaaS vendor with minimal custom code."
  • EventBridge Scheduler (scheduled triggers): Replaces CloudWatch Events scheduled rules. Creates one-time or recurring schedules (cron or rate expressions) that invoke targets — Lambda, Step Functions, SQS, HTTP endpoints — without needing a dedicated event bus. The exam signal: "run this Lambda every night at 2 AM" or "trigger this job every 15 minutes."

EventBridge Pipes (Point-to-Point)

EventBridge Pipes is a newer service that the SAA-C03 and DVA-C02 increasingly include as an answer choice. Know the distinction:

  • EventBridge rules operate on the event bus: events arrive, rules fan them out to multiple targets.
  • EventBridge Pipes connects a single source (SQS, Kinesis, DynamoDB Streams, MSK, Kafka) directly to a single target, with optional filtering and enrichment in between.

The Pipes model: source → [filter] → [enrichment Lambda] → target. The source is always a polling-based service (a queue or stream). No event bus involved.

The exam signal for Pipes: "filter events from an SQS queue or DynamoDB stream and send matching events to a target, with minimal custom code." Before Pipes, you would write a Lambda consumer that manually polled, filtered, and forwarded. Pipes does that without code.


The Decision Framework: Four Questions to Ask

When you see a question that involves AWS workflow or event services, work through these four questions in order:

AWS Step Functions vs EventBridge quick decision tree for SAA-C03 and DVA-C02 exam questions — four-question decision flow with signal words for each outcome

Question 1: Is there a multi-step process where one step depends on the result of the previous step?

If yes, the shape is orchestration → Step Functions. EventBridge cannot enforce "Step 2 only runs if Step 1 succeeds." You could chain EventBridge targets through custom events, but that is exactly the fragile, invisible coordination that Step Functions was built to replace.

Question 2: Does the workflow need retry logic, error compensation, or an audit trail?

If yes, Step Functions. EventBridge rules can invoke retries internally (EventBridge retries failed target invocations with backoff), but those retries are for the delivery of the event — not for a business workflow step that failed for a business reason. Step Functions' Retry and Catch are business-logic constructs that let you model "try the payment up to three times; if all three fail, route to the refund state."

Question 3: Is the trigger an event emitted by an AWS service, a SaaS app, or your application's custom event bus?

If yes and there is no multi-step workflow requirement, the answer is EventBridge. The trigger is a reactive one — something happened, and one or more downstream services should hear about it. This is EventBridge's exact architecture.

Question 4: Does the source need polling (a queue or stream), with filtering, to a single target?

If yes, EventBridge Pipes. The source is one of the polling-based integrations (SQS, Kinesis, DynamoDB Streams, MSK, Kafka). The target is a single downstream service.


Three Exam Scenarios, Solved

Scenario 1: Video Processing Pipeline

A streaming platform uploads videos to S3. Each upload must be transcoded, have thumbnails generated, be stored in a database, and trigger a push notification to mobile subscribers. If transcoding fails, the process must retry automatically up to three times. All steps must complete in sequence, and the operations team must be able to see which step a specific video is currently in.

Correct answer: Step Functions Standard Workflow.

Signal words: "must complete in sequence," "retry automatically," "see which step" (execution history). The operations team visibility requirement rules out Express Workflow (no queryable history). The sequential dependency rules out EventBridge (no enforcement of step order). Standard Workflow stores execution state that the team can inspect.

Scenario 2: Automated Security Response

A security team needs to automatically remediate EC2 instances launched without a required "Environment" tag. When any EC2 instance enters the running state without the tag, a Lambda function should stop the instance and send an alert via SNS. No polling code should be required.

Correct answer: EventBridge rule targeting Lambda.

Signal words: "automatically" (reactive), "EC2 instance enters the running state" (AWS service event — EC2 emits this natively to EventBridge), "no polling code required" (EventBridge receives the event; Lambda does not need to poll). The workflow is one-step ("stop the instance"), not multi-step with dependencies — Step Functions adds unnecessary overhead.

Scenario 3: High-Throughput IoT Telemetry

An IoT platform ingests 80,000 sensor readings per second. Each reading must be validated, transformed, and stored. The transformation logic occasionally needs to be updated without redeploying consumers. The operations team does not need per-reading audit logs.

Correct answer: Step Functions Express Workflow (or, depending on the answer choices, Kinesis + Lambda).

Signal words: "80,000 per second" (Express Workflows handle up to 100,000/sec; Standard would be cost-prohibitive at that volume), "does not need per-reading audit logs" (eliminates Standard Workflow), "validated, transformed, and stored" (multi-step sequence → Step Functions over EventBridge). If EventBridge appears in the answers, its lack of a sequential, dependency-enforced workflow rules it out. Express Workflow wins on throughput and cost.


Putting It Together: When to Combine Both Services

The exam also tests a pattern where Step Functions and EventBridge work together. The architecture: EventBridge delivers the trigger; Step Functions handles the workflow.

Example: An S3 upload event arrives on EventBridge. A rule matches the event and starts a Step Functions state machine (Step Functions is a valid EventBridge rule target). The state machine handles the multi-step processing that follows — with retries, branching, and execution history — while EventBridge handles the reactive, codeless trigger.

The exam signal for this combination: "reduce operational overhead for detecting uploads while maintaining visibility and retry logic over the processing workflow." The "visibility and retry logic" half → Step Functions. The "detecting uploads without polling" half → EventBridge.


Quick Reference: Signal Word Decoder

Signal words in the scenarioService to reach for
"multi-step," "sequential," "pipeline," "workflow"Step Functions
"retry with backoff," "compensate on failure," "error handling"Step Functions
"audit trail," "execution history," "compliance logging per step"Step Functions Standard
"human approval," "wait for callback," "task token"Step Functions Standard
"workflow exceeds 15 minutes" (Lambda limit)Step Functions Standard
"high throughput," "IoT," "100,000 per second"Step Functions Express
"AWS service event" (S3, EC2, RDS, CodePipeline…)EventBridge (event bus)
"SaaS integration," "Zendesk," "Datadog," "Stripe"EventBridge partner sources
"route to multiple targets," "fan-out," "notify multiple services"EventBridge rules
"cron," "schedule," "run every night at 2 AM"EventBridge Scheduler
"filter events from a queue or stream," "point-to-point"EventBridge Pipes
"detect AWS service state change without polling"EventBridge (default bus)

Practice Under Exam Conditions

Service-selection questions like Step Functions vs EventBridge follow predictable patterns — once you've seen 15–20 variations, the signal words map automatically to the right answer. The free SAA-C03 practice questions and DVA-C02 practice set include orchestration and event-driven scenarios in the exact scenario-based format the real exams use.

Not sure whether serverless and event-driven architecture is your current weak spot? Take the free 10-question diagnostic — no signup required — and see your per-domain score in under ten minutes. If Step Functions, EventBridge, or SQS/SNS (covered in depth here) are in your weak domain, you'll know immediately which concepts to prioritize before your exam date.