Oodle · AI Native Engineering
01 / 47
←/→ · T timer · C compact

TIME.

Pens down · debrief

press T to reset
Oodle Finance kickstart + hackathon · 1.5 days

.

One shared operating model for product, engineering, data and finance. Practised on Oodle work, with Claude Enterprise.

Waimakers
The output: a tested prototype, an evidence packet and reusable team improvements.
Your facilitators
Waimakers

Who's in the room.

Abner van den Hout
Abner
van den Hout
AI engineer

Geeks out on agents that ship real work, and on pose-tracking his own backflips.

BackgroundMSc Data Science. Built prediction models for wind turbines and anomaly detection for Dutch trains. Now builds agentic AI platforms, most recently a multi-agent conversational platform for fitness clubs.
These 1.5 daysHands-on builds and codebase work. Co-runs the labs and the hackathon.
Ask me aboutClaude Code, agents, MCP, RAG. Or salto-coach, the Claude skill he built that video-analyzes his water flips and coaches the next attempt.
Maurice Kleine
Maurice
Kleine
Delivery lead

Geeks out on building experimental internet tools that actually work. Life is a series of experiments.

Background14+ years as engineer, product owner, engineering manager and founder. Sold the logistics system he built at university to the company running it on spreadsheets.
These 1.5 daysRuns the room: programme, content and the pre-work interviews.
Ask me aboutAI-native workflows and adoption, getting from prototype to shipped, founding and exits. Or his drum & bass API and LLM puzzle benchmarks.
The sidequest · what a skill can be

Skills can coach backflips.

360° · peak spin 675°/s
360° · tightest tuck 20°
360° · visible flight 4.4s
The sidequest · what one developer ships now

One dev. One month.

GitHub insights for fluncle over one month: 693 merged pull requests, 1,334 commits, 2,328 files changed, 762,042 additions
github.com/mauricekleine/fluncle · fully open source
Get ready · enterprise-safe by default

Bring real work.

Claude Enterprise accessClaude Code, for knowledge work and repository work alike.
An Oodle problem worth reducingA Jira issue, workflow friction, code question or data task. Not a toy brief.
Approved access onlyUse the supplied repositories and systems; never paste secrets or customer data.
A mixed team mindsetProduct frames value, engineering tests feasibility, domain owners verify truth.
Your workshop team · exercises today · hackathon tomorrow

Find your group. Stay together.

Group 01

One

  • Hal White
  • Jennifer Davis
  • Robert Mackay
  • Sudharshini Nagaraj
Group 02

Two

  • Adam Gaventa
  • Jamie Muir
  • Nnamdi Okafor
  • Sourav Sanyal
Group 03

Three

  • Adam Chappell
  • Ambrose Akpoyomare
  • Daniel Haylock
  • Jack Broome
Group 04

Four

  • Josh Ramsay
  • Richard Harris
  • Thomas Hutchins
  • Vivek Sharma
Group 05

Five

  • Hugh Daman
  • Michelle Cleary
  • Nell Jamieson
  • Nic Crane
Group 06

Six

  • Andy Stockbridge
  • Dave Tomlinson
  • Russell Bluck
  • Teagan Black
Group 07

Seven

  • Debayan Roy
  • Nick Costa
  • Poonam Vyas
  • Remy Oudemans
One teamUse this group for the collaborative exercises and the full hackathon. Rotate who drives the agent; everyone owns the result.
Section 01 / 10the shape of the day

.

The wider programme · working programme

Today starts a 12-week build rhythm.

22–23 July

Kickstart

Align, set up and build the first proof.

delivery loop + toolsmini-hackathonplaybook v0.1
After Kickstart

Build cadence

Turn one domain problem into reusable capability.

CLAUDE.md + skill + agentMCP or structured-data workflowpilot reviews + consolidation
10 September

Masterclass

Deepen the full loop and how it is evaluated.

evaluation methodologycaching + batchingagent-building practice
After Masterclass

Sprint cadence

Run real Oodle work through the complete loop.

live ticket or questionrubric + quality reviewssharpen + rehearse proof
15 October

Demo Day

Show the work, peer-review it and decide what advances.

Builder demonstrationsteam-wide challengeadoption decisions
Between sessionsFortnightly virtual coachingLive Teams supportPilot review touchpoints
KickstartIndustrious · 131 Finsbury Pavement · Room A / Lothbury
MasterclassLondon · venue TBC
Demo DayLondon · venue TBC
ThenVirtual wrap + handover · date TBC
These three dates are fixed. Timings, later venues and content emphasis follow what the pilots turn up.
Agenda · one full day + one half day

Learn. Frame. Build.

Day 1 · Wed 22 Jul · align + prepare
01Welcome & baselinewhere Oodle is today09:30
02Setup checkeveryone gets signed in09:50
03The AI-native delivery loopspecify → generate → comprehend & prove10:30
Break11:10
04First hands-onbring Oodle context into a real repo11:25
🍴Lunchoutside, self-organised12:30
05Framing your problemspecify before you build13:15
06Evaluation & trusthow to know what to rely on14:00
Break14:50
07Context & tokensworking efficiently15:05
08Hackathon kick-off + build timeteams, scope lock, Day 2 brief15:50
Close17:00
Day 2 · Thu 23 Jul · half-day hackathon
01Hackathon buildspecify, generate, then prove09:00
02Show & telleach team demos its slice10:45
03Playbook v0.1 + your commitmentsowners, next 30 days11:30
Close12:00
The goal is a proof, not a pitch: a working artefact, its evaluation, honest limits and a named next owner. Teams build in short waves with a checkpoint in the middle.
The program · deliberate compression

The 1.5-day arc.

Day 1 · AMalign + learn
Practicereal repo · safe patterns
Day 1 · PMframe + evaluate
Scope lockproblem · proof · controls
Day 2 · AMbuild, then prove
Evidencedemo · evaluate · decide
Next 30 daysowners + experiments
Every insight becomes a reusable asset: prompt, rule, test, skill or workflow.
Day 1 outputA buildable briefOne user, one pain, one thin slice, one evidence plan and an explicit risk boundary.
Day 2 outputA proof, not a pitchWorking artefact, evaluation results, sources, limitations and a named next owner.
Mission · setup10 min · do this now

Open the tools you’ll use today.

Get the shared working environment ready before we begin.

01
Open Claude Codestart in a blank working folder
02
Connect Atlassian Rovo MCPopen the official setup guide, sign in with your Oodle Atlassian account, and confirm Jira + Confluence access
03
Open the supplied repositoryconfirm Claude Code can see the project
04
Connect the live Q&Aopen the hub and paste its Step 1 command into Claude Code
the huboodle.waimakersmcp.com
10:00
T start / pause
You're done when
Claude Code, Atlassian, the repository and the live Q&A are open and working.
Section 02 / 10where Oodle is today

.

The goal
Build the capability
to work AI-native.
Build your own skillsCapture repeatable ways of working as reusable instructions, tools and checks.
Work with MCPsConnect agents to the systems and context each task needs.
Own the loopFrame intent, guide execution, verify the evidence and accept the result.
Your baseline · 26 of 29 responses · 20 July 2026
Open live Tally results ↗

What you told us.

21use AI most days or all the timefrequency
13have no agentic setup yetsetup
4use an automated evaluation rubricverification
14want help with coding standardsuse cases
20want Jira in the connected workflowsystems
13say the target area is large or oldterrain
Engineering terrain · participant perception
Open live Tally results ↗

The codebases you brought with you.

13/26

Large
or old

Half describe their target area as a large or ageing codebase.

Participant response
9/26

Well
tested

Nine respondents describe the target area as well tested.

Participant response
6/26

Well
documented

Six respondents describe it as well documented.

Participant response
3/26

Newcomer
navigable

Three say a newcomer can navigate the area easily.

Participant response
Tool readiness · who is starting where

Shared appetite. Different starting lines.

21have already used Claudetools
12comfortable or very comfortable in terminalterminal
9have basic terminal confidenceterminal
5prefer to avoid the terminalterminal
9are unsure which rules applyboundaries
6say nothing is off-limitsboundaries
Four depth interviews

AI is already useful. The challenge is making it repeatable.

Already
in the work

Code, Terraform, documentation, extraction and scheduled maintenance are already real use cases.

Interview evidence
useful now · repeatable next
95%
there

Scheduled automation removes work, but can still invent version bumps. Trust still needs checks.

Interview evidence
Usage has a costCopilot CLI includes 100 premium requests per month; active use can hit the limit fast.
Context has a costAtlassian MCP can expose roughly 30 tools and consume 20–40K context tokens before the query.
Reuse is unevenSome teams use skills and agents; many have not yet made them part of the Jira workflow.
Humans own consequenceExplainable, customer and regulated decisions stay with accountable people.
Codebase audit · 14 July, read-only

61 active repos. Zero committed config.

What we scanned311 of Oodle's 691 repos were visible to us. We read the 61 that have been pushed to since 30 June.
What we foundNo CLAUDE.md. No AGENTS.md. No MCP config, no .claude/ or .opencode/ directory, no Oodle-authored skills. We searched every repo of any age and still got zero. The one skills-lock.json came in with an Expo dependency.
What it meansEveryone is using AI. Nobody has committed it. Individual fluency is running well ahead of anything the next person can find in the repo, and that shared layer is what Week 1 puts in: the first CLAUDE.md in a real Oodle repo.
What we build onTwo repos already show the instinct. oodle-ai-customer-summary versions its prompts as files and hot-reloads them. release-notes-automation ships a scheduled job on Copilot CLI. We use those as the examples, not generic ones.
The caveatDefault branches and committed config only. Personal and global setups never show up in an org scan, so "nothing committed" is not the same as "nothing used". We re-run this after the Kickstart to measure what changed.
The goal · how we get there

The goal: turn useful experiments into a repeatable way of working.

The same working loop carries intent into execution, then turns evidence back into reusable practice.
Reuse — decisions, evidence and checks become context for the next task
Frame + connectpeople shape intent and context
Buildagents drive execution
Verifypeople accept the evidence
01Frameagree intent and boundaries
02Connectbring in tools and sources
03Buildexecute the smallest useful loop
04Verifyinspect evidence and accept
Section 03 / 10the delivery loop

.

Background · moving up

The abstraction ladder.

more control · less powermore power · less control →
01Machine & assemblymove every bit by hand
02High-level languagesthe compiler types it out
03Frameworksstop re-solving solved problems
04Natural languagedescribe what you want · same trade, one step up
Background · how we got here

Three eras of AI coding.

you direct every stepagents work on their own →
20% agent
human involvement 80%
01 Tab Autocomplete. You type, it finishes the line.
50% agent
human involvement 50%
02 Agents You direct one agent, step by step. In the loop the whole time.
90% agent
human involvement 10%
03 Cloud agents Hand off and walk away. Many run at once. You brief them and check the result.

Framing: Cursor · Michael Truell, “The third era of AI software development.”

What changes for you

Your time moves to the two ends.

Before today
Specify 15%
Writing code 70%
Review 15%
Toward the goal
Specify 15%
Code 70%
Comprehend 15%
The sandwich · humans on both sides

The agent works in the middle.

SpecifyHumans set the frame: intent, constraints, priorities, taste.
intentconstraintsprioritiestaste?
GenerateThe agent produces the work artefacts.
codethe change
teststhe check
toolsthe reach
MCPsthe context
docsthe memory
patchesthe fix
ComprehendHumans review, polish, synthesise, and decide.
reviewpolishsynthesisedecide

Humans own intent and acceptance. Agents drive execution. Oversight scales with risk.

Coding agents · a new kind of tool

A new tool in the box.

Every tool we had waited for input. You drove it. The coding agent is the first that works on its own.

IDEwrite code
Terminalrun commands
Linterspot mistakes
Debuggerfind bugs
Version controltrack changes
the new tool Coding agent does the work itself: it thinks, acts, and picks what to do next
Shared foundation · what makes the work agentic

A model thinks. A harness lets it act.

Agentic work the outcome you supervise
=
AI model(s) thinks, then picks what to do next
+
Harness tools, files, prompts, and rules: everything that is not the model

A working folder and a repository package different harnesses around the same idea: context, tools, permissions and review shape what the model can safely do.

Coding agents · part one: the model

First, the model itself.

Coding agent
=
AI model(s) thinks, then picks what to do next — let’s see how it reads and remembers
+
Harness

The part you can’t change. To use it well, you need to know how it works.

Crash course · what the model reads

An LLM is autocomplete (sort of).

Guess the next token, add it, guess again.
1 · Words
Is the venue booked?
The model reads your text.
2 · Tokens
Isthevenuebooked?
It cuts the text into small pieces.
3 · Embeddings
0.21-1.40.07·· 1.050.33-0.8·· -0.60.920.14··
Each piece becomes a list of numbers.
4 · They link up
Every token looks back at every earlier one.
Crash course · how it remembers

Tokens & context windows.

Context window
one finite window · quadratic cost
Is _the _venue _book ed ?
Every token attends to the others — that is attention. —— active head
Crash course · what fills the window

Watch the context window fill up.

Hover any line to see what it costs · click a green button to send the next prompt · the real explainer from the Claude Code docs.
Shared foundation · curate context

Long context is not read evenly.

Bury something in the middle of a long prompt and the model is more likely to miss it.
← picked up reliably the middle — where things get missed picked up reliably →
Keep only what is relevant, and put the task after the long source material, not before it. curate · do not dump →
/COMPACTED
Crash course · the failure mode

The dumb zone.

CONTEXT0%

Long session, many files, and the answers went vague. The context window is nearly full. The agent's short-term memory has a hard limit: old details fall off, or blur. This is normal. Now you know its name.

The fix: compact the session, start fresh, or write the state to a file first. Managing long runs is half of what supervision means.
PRESS C · /COMPACT
Coding agents · part two: the harness

Now the other half: the harness.

Coding agent
=
AI model(s) done — we just saw how it reads and remembers
+
Harness tools, files, prompts, and rules — everything that is not the model

The model you can’t change. The harness you build and control. Everything from here is harness.

The harness — the model is one chip on the board; the harness is everything else that makes it useful

Source: Addy Osmani, “Harness engineering” · diagram after Viv Trivedy.

Claude Code architecture — input, knowledge, integration, execution, output, observability and multi-agent layers around the master agent loop

A real harness, annotated by layer · Fareed Khan, “Building Claude Code with Harness Engineering.”

Skills · what & why

What is a skill?

Anthropic · Agent Skills · youtube.com
Skills · recap

The short version.

What it is
A markdown file that teaches Claude something once. After that, Claude uses it on its own, whenever it fits.
When it loads
Only when your request matches its description. Just the name and description sit in context, so it won't fill your context window.
Where it lives
Personal: ~/.claude/skills, follows you across all projects. Project: .claude/skills in the repo, so everyone who clones it gets it.
Crash course · how hard it thinks

Tell it how hard to think.

low
quick, simple tasks
medium
save tokens, a bit less depth
◆ default
high
good for most coding
xhigh
harder problems, more tokens
max
deepest, slowest, costs most
← faster · cheaperdeeper · slower · costs more →
Set the levelType /effort to pick a level. Higher means it thinks more before it acts, and you pay for every thinking token.
One-off deep thinkPut ultrathink in a prompt for one deeper turn. Old words like "think harder" do nothing now.
Crash course · the two dials

Two dials. Engine and thinking time.

/model picks the engine. /effort picks how long it thinks before it acts. That's it.

/model · three tiersFrontier for enduring, hard work. Workhorse as your default. Fast-and-cheap for mechanical bulk. Throwaway work tolerates a cheap model; enduring code earns the frontier one.
/effort · depth, not qualityHigher is not better. It is slower and costs more, and max can overthink a simple ask. Match the dial to the job.
The silent-default trapA low-effort model on a hard task doesn't warn you. It fails confidently. Tracing decision-engine logic wants high effort on a capable model. A rename in the booking repo doesn't.
Labels aren't portable"high" on one model ≠ "high" on another, and switching model can change the supported levels. After /model, check what /effort is actually set to.
Crash course · what agents are bad at

Determinism.

Tweet by Uncle Bob Martin: if anything you want your agents to do is deterministic, such as managing a queue, or sorting a list, create a script for the agent to use. Agents don't do deterministic things well. It's our job to force determinism upon them.
If a task must come out the same way every time, give the agent a script to run. Don't make it guess each time.
Shared foundation · output is a proposal

Fluent is not correct.

> summarise the lending rule and prepare the change
Done. The rule is clear and the change is ready.
source memory · no approved link · no evidence set
! unsupported claim presented with confidence
the control: ground the work in approved sources, start read-only, test the output and keep a named human accountable.
probabilisticapproved.   sources · evaluation · permissions · human decision.
Crash course · control

Permissions & auto mode.

01Ask
every time
02Allow
-list
03Auto
04YOLO
The safe defaultpauses before every action
Trusted throughboring, known-safe commands
Advancedvery helpful, set it up with care
No guardrailsfull speed, where things go wrong
How much freedom you give it. Start tight, loosen as you trust it.
The driving lesson · the tool

Claude Code in the terminal.

A session, your repo, and control over what it may do.

  1. Start it in the repo root. Now it can see your real code.
  2. Permissions. It asks before acting, until you choose to let it run.
  3. Plain language. Point at real files. No magic words.
  4. Esc interrupts. At any time. You are always the one in charge.
~/oodle $ claude
// it reads files. it runs commands. it edits code.
// it asks before it acts.
> explain the booking flow, file by file
Mission · Exercisehands on · 5 min

Block git by default.

You want to run git yourself. Tell Claude Code it can never run git on its own.

add this to your deny list Bash(git *)
01
Open permissionstype /permissions in Claude Code
02
Add the rule to Denypaste the rule above
03
Choose the user scopeso it applies in every project, not just this one
04
Test itask Claude to run a git command · it should refuse
5:00
T start / pause
You're done when
Claude refuses to run any git command on its own.
Mission · Exercise Ahands on · 25 min

Explore the codebase.

Type this first /statusline show my token usage and how full the context window is
01
Explore it broadly"Give me a tour of this codebase: what it does, the main parts, how they fit together." Watch the meter start to climb.
02
Turn it into a report"Write up what you found as a self-contained HTML report I can open." A keepsake of the tour.
03
Go deep on the booking system"Explain the booking system architecture: files, flow, line numbers."
04
Then the user sign-up flow"Explain the user sign up flow: files, flow, line numbers." The meter is getting high now.
05
/clear, then redo step 03Ask about the booking system again, the same way. It starts from scratch. The earlier tour is gone. Compare it to the first answer.
25:00
T start / pause
You're done when
You've filled the window, cleared the session, and seen what it did to Claude's memory of the booking system.
Mission · Exercise Aalso try · debrief

More context management commands.

+
/rewindjump back to an earlier point in the session.
+
/compactshrink the session to a summary and keep going.
+
Write a report with a sub-agent"Write a detailed interactive report of the codebase in HTML. Use a sub-agent." It works in its own window. Yours stays clean.
Debrief
  • After /compact, was the area-A answer as sharp as the first time?
  • What did the meter do across the whole exercise?
  • When would you /compact vs /clear vs hand off to a sub-agent?
Mission · Exercise Bhands on · 12 min

Experiment with models and effort.

Try different model and effort combinations. Get a feel for how answer quality, speed and context use change as you mix them.

01
Switch the modelrun /model and pick a faster, lighter one.
02
Change the effortset /effort low, then high. If your model offers xhigh or max, try one. Notice what the extra thinking costs you.
03
Ask the same thing each timere-run one prompt across combos. Coding: "explain the booking flow, file by file." Non-coding: "draft acceptance criteria for this ticket and list the questions you'd ask first."
Debrief
  • Which combo felt right for quick work? For hard work?
  • Where did a lighter model or lower effort still do the job?
  • When is the heaviest combo actually worth it?
  • Where did more effort just mean more words, not a better answer?
12:00
T start / pause
You're done when
You've tried a few model and effort combos and can say which you'd pick for quick work vs. hard work.
Coffee break
15 min
Stretch your legs. We start again in 15 minutes.
Section 05 / 10specify the problem

.

Work with AI at every step.

Document · decisions become context for the next task §9 later
Specifyyou · the front
Generatethe agent
Comprehendyou · the back
01Aligngrill until you agreethis section
02Planthe spec it will followthis section
03Executethe loop runs & tests§5 ✓
04Reviewread it back, then merge§8 next
Three of the five live up front. Preparing is most of the work.
Mission · Exercise Chands on · 25 min

Implement a small feature.

01
Describe the feature in 1–2 sentencesno more, that's the point
02
Watch the behaviour and the context window
03
Compare your solution with a colleague
Notice
  • Did it follow your coding standards?
  • Did it ask you questions?
  • Would you merge this code, and why?
25:00
T start / pause
You're done when
You have a real opinion on whether you'd merge it, and why.
Mission · Exercise C · part 2hands on · 15 min

Now plan first.

Same feature, fresh start, but let it plan before it writes any code.

turn on plan mode Shift + Tab → until it says plan mode
01
Start cleanrun /rewind or /clear so the old work is gone
02
Turn on plan modepress Shift + Tab until the prompt shows plan mode
03
Ask for the same featureuse the same words as before
04
Compare the two outputswith a plan vs. straight to code
15:00
T start / pause
You're done when
You can name one clear difference between the two runs.
Plan mode · the catch

The problem with plan mode.

PLAN.md
Step 1 · Set up the data layer
Create a migration for the new table
Add indexes on every foreign keyguessed
Write a repository interface and an implementation
Add unit tests for the repository
Update the seed data and the fixtures
Step 2 · Build the service
Define the DTOs and the validation schemas
Implement the service methods
Add error handling, logging and retriesguessed
Wire up the dependency injection
Write the service unit tests and mocks
Step 3 · Wire the API
Add the controller and the routes
Add the request and response models
Add the authorization and rate-limit checks
Add integration tests for each endpoint
Update the OpenAPI spec and the docs
Step 4 · Frontend
Add the form component and its state
Handle the loading, empty and error states
Add client-side validation
Wire it to the new endpoint
Step 5 · Rollout
Add a feature flag and a kill switchguessed
Write the migration runbook
brain
overload
First you just asked. Then you tried plan mode, which was better. But the plan is full of choices it made for you, and it never asked about one. Make it ask first.
Plan mode · why it falls short

Progressive disclosure.

Plan mode
Explore the repo One quick question Writes a big plan You review the wall
You never find out if you actually agree.
What works better
Explore the repo Interview you until you agree Then implement
You build the same picture in your heads first.
You and the agent work together, but you can't see inside each other's heads. Agree on what you're building before any code. After Frederick P. Brooks, The Design of Design: a shared "design concept."
Step 01 · Align · watch first

Grill-me skill

Matt Pocock · "My 'Grill Me' skill went viral" · aihero.dev
Step 01 · Align · the technique

Grill until you agree.

Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree resolving dependencies between decisions one by one.

If a question can be answered by exploring the codebase, explore the codebase instead.

For each question, provide your recommended answer.

find it inparticipant materials · labs day 1
Mission · grill the requirementhands on · 45 min

Agree the brief before you build.

Create a snapshot of what the team currently agrees: the user, outcome, boundaries, evidence of done and open questions. You leave with an agreed brief and three scenarios.

01
WHO · WHAT · WHYuser and situation · required behaviour · intent and outcome
02
NOT · DONEexplicit exclusions and protected areas · pass/fail signals
03
Write three scenarioshappy path · edge case · explicit exclusion in Given/When/Then form
Debrief
  • How much did the context window fill?
  • What changed in the outcome?
  • Quality of the questions?
  • Could you answer simply, or did it need real investigation?
45:00
T start / pause
You're done when
Done when: another team can repeat WHO, WHAT, WHY, NOT, DONE, the proof and red line without asking a question.
🍴
Lunch
45 min
Eat, rest, reset. We start again in 45 minutes.
Section 05 / 10encode it once

.

Same delivery loop · different places where work lives

Claude Code can work wherever the work lives.

01 · working folder

Knowledge
work

Research, requirements, analysis, documents, spreadsheets and PowerPoint or HTML slides.

Outcome · a defensible business artefact
02 · repository

Engineering
work

Explore and change code, tests, architecture, documentation and reviews inside the team harness.

Outcome · verified repository evidence
03 · connected systems

Work across
tools

Reach Jira, Confluence, GitHub and internal tools through MCP, API or CLI.

Outcome · grounded context and action
The interface changes. The working loop stays the same.
Dynamic capability · start from the outcome

Need a capability? Find it—or build it.

What concrete outcome do you need?
01

Name it

Start with the job to be done, not a preferred technology.
Does the capability already exist?
02

Inspect

Reuse or configure what is already available.
What is the smallest missing layer?
03

Build the gap

Add only the capability the outcome actually needs.
Will the team need it again?
04

Share it

Turn repeatable practice into a reusable team capability.
Skill · MCP · API · CLI · script · framework · plugin
There is no fixed routing rule. Use the smallest capability that gets the job done.
Product ↔ engineering · context must survive the hand-off

Intent becomes executable context.

The folder
shapes

  • user + problem
  • evidence
  • options
  • decision

Jira MCP
structures

  • outcome
  • acceptance
  • constraints
  • dependencies

The repo
tests

  • repository fit
  • plan
  • implementation
  • verification
Two sandwiches · one seam must carry the truth

Product intent meets repository reality.

Specify · sources + user
In a folder, Claude Code turns evidence into a decision and requirement.
Comprehend · human challenge
Jira
MCP
acceptance · constraints · links · open questions
Specify · repo plan
In the repo, Claude Code generates inside the engineering harness.
Comprehend · tests + review
Mixed-team handoff · one agreed brief travels; evidence returns

One outcome. One tool. Many hands.

01

Product intent

User, value and thin scope.

02

Working folder

Brief, risks and acceptance.

03

Jira + Confluence

Approved sources and history.

04

Claude Code

Repository map, code and tests.

05

Code evidence

Behaviour, discrepancy and proof.

06

Proof packet

Continue, reshape or stop.

New evidence can change the brief—but the team agrees the change. The agent never changes it silently.
Mission · source-grounded Jira45 min · pairs

Find context. Do not invent it.

Retrieve one pre-vetted Oodle issue and approved Jira/Confluence sources in read-only mode.

01
Retrieve issue + sourcesapproved Jira and Confluence spaces only; preserve links
02
Find at least two gapsmissing user, evidence, rule, constraint, dependency or open question
03
Draft grounded contextlabel every statement sourced, inferred or unresolved
45:00
T start / pause
You're done when
Done when: another pair can trace every claim and name the unresolved questions.
Mission · paired environments45 min · cross-role pairs

Make one artefact you can defend.

Use one agreed brief and approved source pack. Work in parallel, then challenge the handoff.

Working-folder path

  1. Draft requirement and intended outcome
  2. Add risks, acceptance and open questions
  3. Label sourced, inferred and unresolved claims

Repository path

  1. Locate the approved repository slice
  2. Map behaviour, likely files and tests
  3. Capture evidence or discrepancy — no edits
45:00
T start / pause
You're done when
Done when: both paths agree on intent, evidence, one uncertainty and the next human decision.
Mission · Exercisehands on · 30 min

Write a code review skill.

Install the skill-creator skill, then teach it how your team reviews code.

01
Install the skill-creator skillAnthropic's official skill, add it below
02
Make a code review skillhave it scaffold a skill that reviews code changes
03
Add your standards and patternshow your team writes code, and what to flag
04
Run it on real codedoes it catch what you would catch?
/plugins → /skill-creator
30:00
T start / pause
You're done when
Your review skill checks code against your own standards. Ready to use in the next section.
Section 06 / 10permissions, data and consequence

.

Enterprise responsibility · Claude Enterprise is a control, not a waiver

The human owns the boundary.

Oodle
provides

  • approved access
  • data classification
  • business truth
  • accountable owner

The team
controls

  • scope + permissions
  • sources + context
  • tests + review
  • release decision

Claude
assists

  • analysis
  • generation
  • tool use
  • candidate evidence
Risk traffic light · decide before prompting

Match control to consequence.

Green

Public, synthetic or approved low-sensitivity inputs; reversible output; clear human review.

Proceed · keep sources · verify

Amber

Internal context, decisioning logic and rules (contract-protected, allowed), code changes, customer-impacting recommendations or ambiguous permissions.

Pause · reduce scope · add owner + checks

Red

Customer PII, secrets, credentials, autonomous production action, delegating a regulated lending decision, or unsupported compliance claims.

Stop · do not prompt · escalate
Section 07 / 10evidence before confidence

.

Evaluation stack · confidence comes from different kinds of evidence

Do not ask one judge to do every job.

Deterministic

Tests, schemas, exact calculations, linters and fixed expected outputs.

Machine-enforced

Source-grounded

Every business claim traces to approved Jira, documentation, code or data.

Evidence-enforced

AI-assisted

Rubric-based review, adversarial critique and hunting for likely failures.

Human-calibrated

Human-only

Regulatory judgement, customer impact, security risk and production decisions.

Accountable owner
Mission · evaluation stack30 min · define confidence first

Give every check an owner.

Use the agreed brief and the three scenarios you wrote in the grilling lab. Choose only applicable layers, but justify every omission.

01
Deterministic + sourcedtests, schemas, calculations and traceable business claims
02
AI-assistedrubric, adversarial critique and candidate failure discovery
03
Human-onlyregulatory judgement, customer impact, security and production decisions
30:00
T start / pause
You're done when
Done when: every acceptance signal has a method, owner and evidence location, including one failure case.
Review · the techniques

Reviewing AI-generated code.

The agent didn't write it, so it has no bias. Five ways to point one at the diff.

01

Fresh review

A new agent reads the diff in a clean session: one fresh pair of eyes.

02

Many angles

Spawn sub-agents to check from different sides at once. Let Claude split the work.

03

Adversarial

Tell it to break the code: hunt for edge cases and bad input.

04

Security

A focused pass for security holes, like injection or leaked secrets.

05

Self-testing

The agent uses your browser or API to run the code and check it really works.

Mission · Exercise Ahands on · 15 min

Review the code yourself.

Treat it like any other review. Form your own opinion first.

01
Read the diffgo through what the agent changed
02
Write down what you'd changebugs, risks, style · your own list
Debrief
  • What stood out as wrong or risky?
  • How long did it take to feel sure?
15:00
T start / pause
You're done when
You have your own written list of what should change.
Mission · Exercise Bhands on · 15 min

Now let your skill review it.

Run the review skill you just wrote: a fresh agent, no bias, checking your standards.

01
Run your review skillin a fresh context, point it at the branch you just built
02
Compare the two reviewsyours vs. the skill's
Debrief
  • What did the skill flag that you didn't, and the reverse?
  • Did it hold the code to your standards?
  • Where do you agree or disagree? Compare notes with a colleague.
15:00
T start / pause
You're done when
You can point to something your skill caught that you missed, or the reverse.
Mission · Exercise Chands on · 15 min

Run a security review.

We made you a skill for this. Point it at your branch and see what it finds.

01
Run the security-review skilltype /security-review on the code you just built
02
Read every findinggo through what it flags, one by one
03
Judge each onereal risk, or a false alarm?
Debrief
  • Did it catch a real problem?
  • Anything that surprised you?
  • Any findings that were wrong or just noise?
15:00
T start / pause
You're done when
You've judged each finding as a real risk or a false alarm.
Mission · Exercise Dhands on · 20 min

Review from many angles at once.

One agent can run several reviewers at the same time. Each sub-agent checks your code in its own way.

Paste this prompt Review the working tree extensively from four different angles: general code quality, user journey validity, adversarial review, and security review. Spawn a subagent for each. Create a detailed HTML report for me to review the findings.
01
Paste the promptit tells the agent to spawn one sub-agent per angle.
02
Watch them run togetherthe reviews happen at the same time, not one after another.
03
Read each reviewone report per angle: quality, user journey, adversarial, security.
04
Judge and mergekeep the real findings. Drop the noise.
Debrief
  • Which angle caught the most?
  • Did running them at once save time?
  • Any finding one reviewer alone would have missed?
20:00
T start / pause
You're done when
You've read all four reviews and picked the findings worth fixing.
Section 08 / 10Day 1 ends with a buildable brief

.

Compound · self-improving

Self-improving agents.

Run
Capture
Feed back
0laps
Captured decisions
glossary
conventions
guardrails
workflows
Each lap, the next run starts smarter.
Compound · the project's memory

What is CLAUDE.md?

Anthropic · Claude Code · youtube.com
CLAUDE.md · recap

The short version.

What it is
A markdown file in your project root. Claude reads it every time you start a session. Its text is added to your prompt: an onboarding note for your codebase.
How to make one
Run /init and Claude drafts one from your code. When you correct Claude, ask it to save that rule to memory.
Where it lives
Project: in the repo root, shared with your team. Personal: in your config folder, just for you, across every project.
CLAUDE.md · what goes wrong

How it turns into a mess.

×The ball of mud
You add one rule each time Claude slips. Months later the file is huge, and the rules start to fight each other.
×Auto-generated bloat
Letting /init fill the file with "useful for most cases" rules. Most of it isn't needed today. Trim it down.
×Past the budget
Every word loads on every request. A model only follows about 150–200 rules well. A bloated file leaves less room for the real task.
×Stale paths
Don't write where files live. Paths move, and Claude will trust the old one. Say what the project does instead.
Keep the root file tiny: one line on what the project is, the package manager, the build command. Push the rest into separate files, and load them only when needed. Based on Matt Pocock's writing on AGENTS.md / CLAUDE.md.
Mission · Exercisehands on · 20 min

Make your CLAUDE.md, then trim it.

Let Claude draft one, read it closely, and cut it down to what matters.

01
Run /initlet Claude draft a CLAUDE.md from your codebase
02
Read it carefullyevery line: is it true, and does every task need it?
03
Trim it downcut bloat, drop file paths, fix anything stale
/init
20:00
T start / pause
You're done when
Your CLAUDE.md is short and true: what the project is, the package manager, the build command. Nothing stale, no bloat.
Compound · see it in action

The grill-with-docs skill.

Matt Pocock · youtube.com
Compound · a skill that does the loop

The grill-with-docs skill.

It grills you
"You said 'account'. Do you mean the Customer or the User?"
"Your glossary defines 'cancellation' as X, but you seem to mean Y. Which is it?"
"Your code cancels whole orders, but you just said partial. Which is right?"
One hard question at a time. It tests your plan against your code and against the words your team already agreed on.
It writes the answers down
CONTEXT.md a glossary: the exact words your team agrees on, and nothing else
docs/adr/ decision records: why you chose this, kept for the next person
Captured the moment a decision is made, not batched up for later.
This is the loop from the last slide, made real: every session leaves the docs a bit clearer, so the next run starts smarter. Based on Matt Pocock's grill-with-docs skill.
Mission · Exercisehands on · 15 min

Get the grill-with-docs skill.

Download it, add it to a project, and read it closely: what makes it good?

01
Download the skillget the grill-with-docs folder from the link below
02
Add it to a projectdrop the folder in .claude/skills
03
Read SKILL.md closelyhow does it grill, and how does it write things down?
participant materials · reusable context lab
15:00
T start / pause
What makes it good?
  • Why ask one question at a time, and always give a recommended answer?
  • Why write CONTEXT.md and ADRs as you go, not at the end?
  • When does it say to skip an ADR?
Done when
Done when: your team identifies one pattern to reuse and one safeguard to add.
The model, at scale

The factory is many sandwiches at once.

Many agents work at the same time. But only one person reviews. That is the slow part.

Agents spread out. Many tasks generate at the same time.
Human review is the one slow step. It sets how fast the factory really goes.
Run more agents than you can read and you do not go faster. You just approve without checking.
Reviewone human · one slow step
The queue is growing. The bottleneck is you.
Mission · make it reusable20 min · choose your path

Teach the next run.

Capture one repeated decision so tomorrow’s work starts with Oodle context instead of a blank prompt.

ENGINEERING

CLAUDE.md

Add one repository rule, verification command and protected boundary.

PRODUCT · DATA · RISK

Knowledge-work Skill

Package one repeatable brief, source checklist and review gate.

20:00
T start / pause
You're done when
A teammate can reuse the instruction, find its sources and see exactly where human approval remains.
Readiness gate · no team starts Day 2 with a vague pitch

Five green checks.

Agreed brief WHO · WHAT · WHY · NOT · DONE draftedWith happy path, edge case and explicit exclusion.
Proof evaluation set, sources and decision rule namedThe team knows what continue, reshape and stop mean.
Boundary traffic light, data, permissions and red line approvedAmber has a reviewer and controls; red does not proceed.
Team product, engineering, domain, risk and proof lenses coveredPeople may hold two roles; no accountability is missing.
Access approved systems and synthetic fallback testedRepository, working folder, Jira/Confluence MCP and normal test command.
One address · everything for tomorrow lives here

Find it on the hub.

Hackathon pack brief · inspiration pack · three startersEight build ideas, gentlest on-ramp first. Three with the brief pre-filled. Grab one, skip the blank page.
Skills starter guide which skill for whatTerminal on-ramp per skill, sources, and the 90-second install demo. Read the SKILL.md first.
Labs & templates Day-1 labs · requirement · rubric · proof packet · riskEverything you filled in today, blank and reusable tomorrow.
Pull it into Claude Code the hub is an MCP serverConnect once and say "grab the extraction starter and grill me". No copy-paste.
the huboodle.waimakersmcp.com
Section 09 / 10day two · the proof gets built

.

Flagship challenge · one bounded full-cycle audit

Trace a rule from intent to evidence.

01Business ruleowner + approved source
02Jira contextdecision + acceptance
03Approved code sliceread-only, bounded
04Discrepancy evidencesource + test + limit
05Improvement ticketowner + human decision
business rule → Jira context → approved code sliceOne logic slice. Read-only first. No scored lending decision delegated to AI.
Mixed teams · everyone builds; accountability stays explicit

Five lenses. One proof.

01

Product
lead

Owns the user, value, thin scope and evidence narrative.

02

Engineering
lead

Owns repository fit, implementation, tests and technical review.

03

Domain
verifier

Owns business truth, data meaning and acceptance.

04

Risk
lead

Owns permissions, traffic light, sources and limitations.

05

Proof
lead

Owns captured evidence, reusable asset and five-minute demo.

FormationEven mix of technical and non-technical: at least one engineer per team, two where possible. And one named captain who breaks ties, cuts scope and owns the demo clock.
Same seven teams · no reshuffle on Day 2

Your group is your build team.

Group 01

One

  • Hal White
  • Jennifer Davis
  • Robert Mackay
  • Sudharshini Nagaraj
Group 02

Two

  • Adam Gaventa
  • Jamie Muir
  • Nnamdi Okafor
  • Sourav Sanyal
Group 03

Three

  • Adam Chappell
  • Ambrose Akpoyomare
  • Daniel Haylock
  • Jack Broome
Group 04

Four

  • Josh Ramsay
  • Richard Harris
  • Thomas Hutchins
  • Vivek Sharma
Group 05

Five

  • Hugh Daman
  • Michelle Cleary
  • Nell Jamieson
  • Nic Crane
Group 06

Six

  • Andy Stockbridge
  • Dave Tomlinson
  • Russell Bluck
  • Teagan Black
Group 07

Seven

  • Debayan Roy
  • Nick Costa
  • Poonam Vyas
  • Remy Oudemans
First moveChoose a captain, then assign the five lenses on the previous slide. One person may own more than one lens.
Mission · specify45 min · mixed teams

Agree the brief.

Product owns the user and value; engineering owns feasibility and verification; the domain owner owns truth.

01
WHO · WHAT · WHYuser and situation · required behaviour · intent and outcome
02
NOT · DONEprotected areas and exclusions · pass/fail acceptance signals
03
Grill and agreehappy path, edge case, exclusion, source, traffic light and non-goals
45:00
T start / pause
You're done when
Done when: the five-field brief, examples, source, proof and red line are agreed by the team.
Mission · generate75 min · midpoint checkpoint

Build the agreed slice.

A working folder owns business artefacts; the repo owns repository work. Neither may silently change the agreed brief.

01
Produce the thinnest pathone approved input → one transformation → one useful output
02
Checkpoint at 35 minutesshow a working path or cut scope; record every brief question
03
Capture evidence as you workcommands, sources, input/output, tests and agent decisions
75:00
T start / pause
You're done when
Done when: the agreed thin slice works end to end and its representative input/output is captured.
Mission · comprehend & prove45 min · evidence gate

Prove what you trust.

No new features. Read the result, challenge it and complete the proof packet.

01
Run the evaluation stackdeterministic · source-grounded · AI-assisted · human-only
02
Show one failureknown limit, rejected output or recovery path · with evidence
03
Package the learningreusable rule, test, skill, Jira flow or playbook plus named owner
45:00
T start / pause
You're done when
Done when: the proof packet supports continue, reshape or stop without relying on the pitch.
Final handoff · turn the build into a crisp shared idea

Put your proof on the idea wall.

TitleA short name people will remember.
ProblemWho struggles today, and with what?
IdeaThe thinnest useful workflow you propose.
ProofWhat you built, tested or learned.
submit_group_idea

Tell your agent your group number and these four fields. Submit again to update your group's card.

Presenter wall · oodle-present.waimakersmcp.com/ideas
No pitch theatreThe card sets the claim. Your five-minute demo supplies the evidence, the failure and the decision.
Section 10 / 10what stays with you

.

Proof packet · the demo is one part of the evidence

Show the work. Show the limits.

01

User +
claim

Who improved, what changed and why it matters.

02

Working
slice

Representative input, live output and before/after.

03

Evaluation
results

Commands, rubric, fixed set, scores and failures.

04

Sources +
limits

Traceability, assumptions, red line and known failure.

05

Owner +
decision

Reusable asset, next experiment and continue/reshape/stop.

Wrap up · what did we learn

What did we learn?

Specify a good spec is most of a good result.
Supervise long runs are managed, not just survived.
Comprehend review is where quality is set, not a quick approval.
Compound a shared oodle-context skill, CLAUDE.md and Skills make every next run start with Oodle context.
From Monday I will · make the next move explicit

Turn the workshop into work.

Write one commitment that a colleague can observe and review.

Real workflow name the product, engineering, data or risk task you will run.
Four-week use case keep it bounded enough to evaluate, improve or stop.
Human gate + evidence name who approves and where the proof will live.
Playbook v0.1 · the cohort’s first operating system

Capture what deserves to compound.

Pattern the prompt, Skill or CLAUDE.md instruction that worked.
Safeguard the source, evaluation, permission or human gate it needs.
Owner the person who keeps it accurate, useful and approved.
Next experiment one measurable use case to run for four weeks.
How context compounds · your strategy's three layers, made real

One shared brain. Three layers.

oodle-context skill one skill everyone installs: who Oodle is, Consumer Duty, hard rules. SteerCo owns it.
Project config per-function skills, MCPs and CLAUDE.md / AGENTS.md. Champions own it.
Builder skills reusable prompts and skills, where Oodle's AI IP accumulates. Builder champions own it.
Today's outcome a shared team repo with these configured, or a clear vision plus 1–2 named owners to build it.
Next up a shared compliance agent: a pre-commit check or PR scanner that flags Consumer Duty issues for a human. Built on shared context; never deciding for you.
Bonus · §10if there's time

.

BONUS · autonomy with a target

/goal & workflows.

goal · checkable pass/fail docker image
0% size reduced
target −20%
01 build 02 measure 03 trim 04 re-check
Give it a number it can check — it loops until it gets there. You orchestrate, not type.
Mission · Exercisehands on · 25 min

Set a goal. Watch it run.

Write down a measurable goal, then use workflows to chase it.

01
Write one measurable goale.g. "decrease Docker image size by 20%"
02
Use /goal and a workflow
03
Watch it work, and step in when it's wrong
25:00
T start / pause
You're done when
The number moved, and you understand what it changed to get there.
Mission · Bonus exercisehands on · 35 min

Build with the whole machine.

Pick a feature that touches many concepts, and use everything you've built today.

01
Use your updated "grill with docs" skill
02
Use your agent.md and your docs
03
Capture decisions as you go
Debrief
  • What important decisions got documented?
  • Did the LLM follow your agent.md?
35:00
T start / pause
You're done when
Your repo is smarter than it was this morning, and you can show where.
Mission · Bonus exercisehands on · 20 min

Let an agent test it like a user.

Give an agent a browser and let it click through your feature.

01
Install the Chrome MCPthis lets the agent control a browser
02
Spawn an agent with browser accesspoint it at the running app
03
Have it click through the featurelike a real user would
Debrief
  • Did it find anything that broke in the real UI?
  • What could it not test on its own?
20:00
T start / pause
You're done when
An agent has clicked through your feature in a real browser.
Close · the workshop, done



Every team moved through Specify → Generate → Comprehend & Prove on Oodle work, and left evidence, limitations and a reusable improvement behind.

The next step is not more prompting. It is repeating the controlled loop on valuable work, measuring the result and improving the harness. Oodle remains accountable for every decision.

Specify Generate Comprehend + prove
materials · prompts · answersoodle.waimakersmcp.com