The clinical research team had spent weeks on the Event Model. The patient eligibility flow sprawled across multiple swimlanes—Clinical Research Coordinators entering data, an automated assessment engine, and Principal Investigators making final calls. Orange sticky notes for events lined the top timeline. Blue cards for commands sat in whiteboard arrays. Green read models showed what each actor saw when they made decisions.
It was a beautiful map. And then someone had to translate it into tests.
Three developers, two QA engineers, and a business analyst spent three days writing Given/When/Then scenarios by hand. Matching sticky notes to Gherkin syntax. Cross-referencing arrows to assertions. Manually ensuring every slice represented in the visual model had corresponding test coverage.
By day three, one of the developers pulled up the Event Model on her laptop and held it next to the test file. “We’re literally copying this diagram into text,” she said. “Verbatim. Why are we doing this by hand?”
That question—asked in frustration—points to the next evolution of executable specifications. The scenario is the specification, yes. But what if the specification could generate the scenario?
The Translation Tax
In the previous post, I argued that EventModeling collapses specification and testing into one artifact. The Given/When/Then format becomes the single source of truth that both describes system behavior and verifies it.
But here’s the gap I didn’t address: someone still has to bridge from visual model to executable text.
That bridge has a cost. Every slice in your Event Model becomes a test scenario. The orange events (past facts) become the Given. The blue command becomes the When. The resulting event—or error—becomes the Then. Each slice is a complete behavior specification.
The translation isn’t complex—it’s mechanical. Tedious. Error-prone in exactly the way you’d expect from repetitive manual work.
I’ve watched teams:
- Miss slices entirely (that sticky on the diagram nobody noticed until production)
- Mix up preconditions (using PatientScreened when they meant PatientEnrolled)
- Create redundant scenarios (testing the same slice three times because it appeared in different swimlanes)
- Drift from the model over time (tests updated, diagram untouched, connection forgotten)
The separation between visual model and executable test reintroduces the very problem EventModeling solves. Two artifacts that say the same thing, maintained separately, eventually saying different things.
There’s a better way.
What the Model Already Knows
Look closely at an Event Model slice. Each sticky note contains everything needed to generate the test scaffold:
Orange events (facts that occurred) define the Given. Each event has a type name and payload schema. Zero to many events describe what must exist as historical fact before the behavior triggers.
Blue commands (imperative, issued by an actor) define the When. Each command has a name and input. The actor provides this input—sometimes explicitly through a form field, sometimes implicitly through the UI action they trigger (like selecting a patient record).
The resulting event or error defines the Then. One slice, one outcome: either a new past-tense event is emitted, or an error is raised because business rules prevent the command.
Green read models show what the actor sees when they make their decision. These inform the scenario’s context but don’t directly appear in it—the command’s timestamp captures when the actor acted on what they saw.
Swimlanes (rows) define which actor issues the command. Tests naturally group by the actor who triggers the behavior.
Here’s the critical insight: In Event Modeling, there are no conditional branches within a slice. A branch means two separate slices.
Patient over 75 with elevated enzymes? That’s one slice with its own Given/When/Then. Patient under 75? That’s a different slice. Patient exactly 75? Another slice. The “branching” happens at the model level—different business scenarios modeled as different slices—not at code level with if-statements in a single scenario.
The model is the specification in visual form. The Given/When/Then format is just a text serialization of that same structure—one scenario per slice.
From Diagram to Tests: How It Works
Let’s walk through a concrete example showing the mechanics.
Event Model Slice S1: Customer issues a command to confirm payment with available funds.
SWIMLANE: Customer (sees OrderSummary green read model)
[OrderPlaced] ───────────────────────────────┐
▼
[ConfirmPayment]
│
▼
[PaymentConfirmed]
The algorithm walks this slice:
- Given: One orange event (OrderPlaced) that must exist as historical fact
- When: Blue command (ConfirmPayment) issued by Customer
- Then: One orange event (PaymentConfirmed)
Generated test scaffold:
Feature: Payment Confirmation
Scenario: S1 - Payment succeeds when order placed
Given OrderPlaced has occurred
| orderId | amount | customerId |
| {{id}} | 150.00 | {{cid}} |
When Customer issues ConfirmPayment
| orderId | paymentMethod |
| {{id}} | credit_card |
Then PaymentConfirmed occurs
| orderId | confirmedAt | paymentMethod |
| {{id}} | {{now}} | credit_card |
Notice:
- Orange OrderPlaced became the Given precondition
- Blue ConfirmPayment became the When action, with the actor (Customer) issuing it
- Orange PaymentConfirmed became the Then outcome
- Green OrderSummary is the read model the Customer sees, informing their decision—but the test captures the command they issue as a result
The customer doesn’t provide all values explicitly. The orderId comes from the OrderSummary they were viewing (implicit from their UI context). The paymentMethod they select explicitly. The timestamp of their action is implicit.
Handling Multiple Scenarios: Separate Slices
What about the “insufficient funds” case? In Event Modeling, this isn’t a branch—it’s a separate slice.
Event Model Slice S2: Customer attempts to confirm payment but fails due to insufficient funds.
SWIMLANE: Customer (sees OrderSummary and AccountBalance)
[OrderPlaced] ───────────────────────────────┐
▼
[ConfirmPayment]
│
▼
[error: InsufficientFunds]
Generated test scaffold:
Scenario: S2 - Payment fails when insufficient funds
Given OrderPlaced has occurred
| orderId | amount | customerId |
| {{id}} | 150.00 | {{cid}} |
When Customer issues ConfirmPayment
| orderId | paymentMethod |
| {{id}} | credit_card |
Then InsufficientFunds error occurs
| orderId | requestedAmount | availableBalance |
| {{id}} | 150.00 | 50.00 |
The Customer actor sees different green read models in this scenario—the OrderSummary (which patient they’re working with) AND the AccountBalance (which informs them they lack funds). The command itself is identical. The past events are identical. But the outcome differs because this is a different business scenario: one where funds are insufficient.
The model makes this distinction visible. Two slices, two scenarios, zero conditional branches.
Complex Preconditions: Zero to Many Events
Back to the clinical trial example from the previous post—the patient eligibility rule.
Event Model Slice S3: Screening System assesses eligibility for a patient who is over 75 with elevated liver enzymes, requiring PI review.
SWIMLANE: Screening System (sees PatientScreeningData)
[PatientEnrolled] ───────────────────────────┐
▼
[ScreeningCompleted] ──────────────────┐ [AssessEligibility]
- age: 78 │ │
- liverEnzymes: "elevated" │ ▼
- creatinine: "normal" │ [PIReviewRequired]
- ... │
▼
[EscalationFlagged]
Generated test scaffold:
Feature: Patient Eligibility Assessment
Scenario: S3 - Patient over 75 with elevated enzymes escalates to PI
Given PatientEnrolled has occurred
| patientId | enrollmentDate | studyId |
| p-789 | 2026-01-10 | S-2026-A|
And ScreeningCompleted has occurred
| patientId | age | liverEnzymes | creatinine | bmi |
| p-789 | 78 | elevated | normal | 24.5 |
When Screening System issues AssessEligibility
| patientId | screeningDate |
| p-789 | 2026-01-15 |
Then PIReviewRequired occurs
| patientId | reason | triggeredBy |
| p-789 | Age over 75, elevated enzymes | EligibilityAssessment |
And EscalationFlagged occurs
| escalationId | patientId | priority |
| esc-456 | p-789 | high |
Wait—you say—”And” in the Then? That’s two events!
Yes. If the business outcome requires multiple facts to be recorded atomically (same transaction, same consistency boundary), they appear together in the Then. The generator knows this because the slice shows multiple orange events downstream of the command. They’re not branches—they’re a single business outcome.
The Given has two events (PatientEnrolled preceded ScreeningCompleted). Zero to many events describe the historical state. The actor—the Screening System—sees the PatientScreeningData read model and issues AssessEligibility. The decision inputs (age 78, liverEnzymes elevated) come from the past events and the read model context.
The Coverage Promise
Here’s where generated tests fundamentally change the game.
When a human writes tests from a diagram, coverage is optional. You hope the QA engineer noticed that sticky note. You assume the developer understood that edge case. You trust the team updated tests when the model changed.
When tests generate from the model, coverage is structural. Every slice is a scenario. Every orange event in a slice becomes part of a Given or Then. Every blue command becomes a When. The only way to have untested behavior is to have unmodeled behavior.
This inverts the usual discovery pattern:
- Before: Teams discover gaps in testing during QA review or (worse) production incidents
- After: Teams discover gaps in the model during generation—”Wait, we have no slice for exactly age 75″—and add them as explicit slices
The Event Model becomes a coverage checker. Gaps are visible. Missing scenarios stand out. The act of modeling and the act of specifying converge.
Multiple Given Events: Building State Over Time
Let’s see a slice with more complex preconditions—zero to many events in action.
Event Model Slice S4: CRC attempts to reschedule a patient appointment where the patient has already rescheduled twice before.
SWIMLANE: CRC (sees PatientSchedule and RescheduleHistory)
[PatientEnrolled] ───────────────────────────┐
▼
[AppointmentScheduled] ────────────────┐ [RescheduleAppointment]
- appointmentId: a-100 │ │
- date: 2026-03-15 │ ▼
- status: initially_scheduled │ [error: RescheduleLimitExceeded]
│
[AppointmentRescheduled] ──────┐ │
- appointmentId: a-100 │ │
- previousDate: 2026-03-15 │ │
- newDate: 2026-03-22 │ │
│
[AppointmentRescheduled] ──────┘ │
- appointmentId: a-100 │
- previousDate: 2026-03-22 │
- newDate: 2026-04-05
Generated test scaffold:
Scenario: S4 - Reschedule fails when patient exceeded limit
Given PatientEnrolled has occurred
| patientId | enrollmentDate |
| p-555 | 2026-01-01 |
And AppointmentScheduled has occurred
| appointmentId | patientId | date | status |
| a-100 | p-555 | 2026-03-15 | initially_scheduled|
And AppointmentRescheduled has occurred
| appointmentId | previousDate | newDate |
| a-100 | 2026-03-15 | 2026-03-22 |
And AppointmentRescheduled has occurred
| appointmentId | previousDate | newDate |
| a-100 | 2026-03-22 | 2026-04-05 |
When CRC issues RescheduleAppointment
| appointmentId | requestedNewDate |
| a-100 | 2026-04-12 |
Then RescheduleLimitExceeded error occurs
| appointmentId | currentRescheduleCount | maxAllowed |
| a-100 | 2 | 2 |
Four past events establish the state. The CRC sees PatientSchedule (the appointment details) and RescheduleHistory (the count) through green read models. Their command includes the appointmentId (likely from the schedule view, implicit context) and the requested new date (explicit input). The outcome is an error—business rules prevent this command given the historical state.
Zero to many events. One command. One outcome.
The Regeneration Ritual
Generated tests sound great at day zero. What about month six?
This is where the approach proves itself. The lifecycle becomes:
- Business discovers new requirement (compassionate use protocol needs different rules)
- Team adds new slice to Event Model (separate slice with different preconditions or different outcome)
- Generator creates new scenario (the compassionate use slice appears automatically)
- Team fills in concrete values (clinical researcher provides test data)
- Developer implements until tests pass
The model is the master. The tests follow. When the model changes, regeneration shows exactly what changed in the specification. The diff between old and new generated tests is the specification change, expressed in executable form.
This matters enormously for governance and audit (topics from earlier posts in this series). Regulators don’t want to hear that someone “updated the test suite.” They want traceability: requirement → model → test → result. Generated tests provide mechanical traceability. A human decided to change the model. Everything downstream flowed mechanically.
Tooling the Pipeline
You don’t need sophisticated tooling to start. The basic pipeline:
Input: Event Model in any structured form
- JSON export from digital whiteboard (Miro, FigJam)
- Simple YAML/JSON describing slices (events, command, outcomes)
- Even a spreadsheet with slice ID, Given events, When command, Then outcome
Generator: Template engine
- For each slice, emit Given/When/Then template
- Populate from event schemas
- Add placeholders for concrete values
Output: Executable scenarios
- Gherkin for Cucumber/SpecFlow
- Jest/Mocha test skeletons
- Custom format matching your testing framework
CI Integration:
- Event Model changes trigger regeneration
- Generated tests committed alongside manual value-fill
- Test failures indicate model drift or implementation gaps
The sophistication can grow:
- Event schema validation (ensure events have required fields)
- Consistency boundary checking (verify atomic outcomes)
- Traceability reports (link each test back to slice coordinates)
But the core concept—slice → scenario → executable—is simple enough to prototype in an afternoon.
The Psychological Shift
Remember the psychological benefit from the previous post? The confidence that specification and test are the same thing—no drift, no ambiguity, no wondering what’s true?
Generated tests multiply that confidence.
When tests are handwritten, there’s always a nagging question: Did someone miss something? Is the test suite complete? Does it actually match the model we agreed on?
When tests are generated, those questions become mechanical. The suite is the model’s expression. The only completeness question is: Is the model complete? And that’s a business conversation, not a code review.
Developers stop asking “did I test the right thing?” They know. The model told them exactly what slices exist.
Business stakeholders stop asking “how do I know the tests match what we agreed?” They look at the model. The tests came from it. Directly. Provably.
Auditors stop asking “show me traceability from requirements to tests.” The generation log is the traceability. Requirement expressed in diagram. Diagram compiled to test. Test executed in CI.
From Specification to Generator
Let’s bring this back to the central thesis of the series.
EventModeling creates a shared language between business and technology. The visual model bridges disciplines. The scenario format makes behavior executable.
Adding generation completes the loop: business intent → visual model → executable specification → running system. Each step is traceable. Each transformation is mechanical or structured.
The speculative document from the original whiteboard becomes the test harness for the production system. The sticky notes from the Event Modeling session become the assertions in your CI pipeline. The conversation between clinical researcher and developer—captured visually, preserved mechanically, verified continuously.
This is specification as code, taken seriously. Not requirements in a document we hope someone maintains. Not tests someone wrote based on their interpretation. But the system specification, expressed visually, generated mechanically, verified automatically.
The Event Model isn’t documentation. It’s source code for behavior. And like all source code, it can be compiled.
What This Makes Possible
Some implications worth exploring in your own practice:
Living documentation that actually lives. The Event Model is the only artifact you maintain by hand. Tests regenerate automatically. The visual and executable stay synchronized by construction.
Test maintenance as model discussion. When requirements change, the conversation happens at the whiteboard. Everyone sees whether this is a new slice or a change to an existing one. The new tests appear automatically, diff cleared, ready for domain values.
Coverage by structure. You can’t forget to test a slice if every slice generates a scenario. You can’t miss an edge case if the model makes it visible as a separate slice. Coverage becomes a property of modeling quality, not testing diligence.
No conditional branches to miss. Because Event Modeling makes every scenario explicit as its own slice, there’s no hidden if-statement logic. The “branch” is right there on the board—two slices side by side.
Audit without archaeology. The path from regulatory requirement to executed test is mechanical. Requirement appears in model. Model generates test. Test runs in CI. No missing links. No “we think this covers it.”
Onboarding as exploration. New team members read the Event Model to understand the system. They run the generated tests to see behavior in action. They add slices when they find gaps. The cycle reinforces itself.
The Question For Your Team
The developer in the opening story asked the right question: Why are we doing this by hand?
The answer, usually, is: Because we always have. Because the tools don’t support generation. Because it seemed too hard to automate.
None of those answers hold up once you see the mechanics. The Event Model is structured data. Tests are structured text. Template generation is standard.
If your team is already doing EventModeling—already sketching orange events and blue commands on a whiteboard—then you’re one generator away from executable specifications that never drift from the visual model.
If you’re not doing EventModeling yet, the existence of generation is one more reason to start. The upfront modeling work—already valuable for alignment and shared understanding—pays compounding dividends when it becomes your test suite’s foundation.
The specification becomes the test when you can read them both and see they’re the same story. The specification generates the test when your tooling compiles one into the other.
Both tell the same story that was told at the beginning. Both can finally be believed.