Proddojo

Real repositories three modes graded against the real fix

Tutorials teach you to code.
This teaches you to ship.

Production is two hundred files you did not write, a deploy that fails for reasons the error does not name, and logs as the only clue. Proddojo drops you into that, grades your fix, and shows the commit that fixed it for real.

shape: per-seat subscription, individuals and teams grows from: mockterview, azkaban, envdiff
proddojo challenge 47 · the silent failure debuggingobservability
scenario POST /checkout returns 200. Orders are not in the database.
No errors in the logs. Customers are complaining.
environment api · checkout · inventory · postgres running
docker logs, psql, git blame
source open-source project under a permissive licence, credited on the challenge
$ docker logs checkout | tail -4
14:23:01 INFO  Publishing order.created
14:23:01 WARN  Queue connection timeout, retrying
14:23:07 WARN  Queue connection timeout, giving up
14:23:07 INFO  POST /checkout completed 200
time taken00:41
tests passed12/12
graded7/10
  • root cause named
  • fix scoped to the publisher
  • regression test added
maintainer's fix a1c9e02 "fail the request when the publish gives up"
Example output. The last line is the point of the exercise: the fix that shipped, beside yours.
challenge=47
the silent failure: orders return 200 and never reach the database
median_to_root_cause=38m
across everyone who has solved it
rubric=7/10
graded by a second model against the maintainer's actual fix

Readings from the last run.

The gap

Tutorials end. Production begins.

Debugging under uncertainty is the skill that closes the distance, and it is the one nothing on the market teaches on purpose.

What tutorials teach

  • Todo apps
  • Clean starter projects
  • One file at a time
  • "It works on my machine"
  • Perfect happy paths

What production requires

  • 200-file codebases
  • Legacy code nobody documented
  • Services that talk to services
  • Logs as your only clue
  • Edge cases that break everything

Proddojo is built entirely out of what production requires.

A self-taught developer or a bootcamp graduate can build an application from a clean start. The first production codebase is nothing like that: unfamiliar structure, a build that fails for reasons the error does not name, three services talking to each other, and a bug report that says orders are missing and nothing else.

Interview platforms test algorithms. Courses test syntax. What hiring managers test, in take-home tasks and first weeks, is whether someone can read logs, follow a stack trace through code they did not write, and change it without breaking the rest.

The bugs already exist. Open-source projects fix thousands of them a year, each with the report, the discussion and the commit that closed it. Packaged as challenges, they are the closest thing to a first week that can be practised.

The form

A brief, a running environment, a graded fix, and the real one beside it

Four movements, in order, every time. The environment is live before the brief is read, and nothing is multiple choice.

01
brief

Read the brief

A scenario in the words a customer or a colleague would use, and nothing else.

02
investigate

Investigate

A real environment in the browser: the services running, docker logs, psql and git blame available. No multiple choice. Hints on request, from a model that has read the fix and will not reveal it.

03
submit

Fix and submit

The change runs against the project's own tests in an isolated container, with azkaban doing the isolation.

04
review

Learn

A second model grades the fix against a rubric built from the maintainer's actual commit, the way Mockterview grades an interview with a model that never conducted it. The real commit is shown beside yours.

Three modes

Find it, build it, or read the diff that someone else wrote

The same environment and the same grader, pointed at the three things a first month actually asks for.

mode 01

Debug

A report that names a symptom. The cause is somewhere in a codebase nobody in the room wrote.

  • Orders return 200 and never reach the database
  • Memory climbs for six hours, then the pod restarts
  • The migration locks the table it was meant to widen
graded onroot cause, scope of the fix, regression test
mode 02

Build

A change added to a repository that already has opinions, without breaking what depends on it.

  • Paginate an endpoint that returns the whole catalogue
  • Make the retry idempotent without changing the public API
  • Backfill a column on a live table with no write lock
graded onbehaviour, blast radius, tests written
mode 03

Review

A diff to read the way a colleague would, then approve or block it, with the reasons written down.

  • A pull request that passes every test and drops a tenant filter
  • The diff behind last Friday's incident, before it shipped
  • A cache put in front of a read that goes stale
graded onwhat was caught, what was missed, how it was said

Every challenge in all three modes is drawn from an open-source project under a permissive licence, and the project, the issue and the fixing commit are credited on the challenge itself.

Progression

Eight skills, one rank each, earned only against real repositories

A rank moves when a challenge is solved without a hint and the grader agrees with the maintainer. It does not move for time served.

Ientering IIworking IIIfluent IVtrusted
Debuggingrank III fluent

11 challenges, 2 with a hint

Observabilityrank III fluent

7 challenges, logs and traces

Databasesrank II working

5 challenges, locks and migrations

APIsrank III fluent

9 challenges, contracts and retries

CI/CDrank II working

4 challenges, red pipelines

Securityrank I entering

2 challenges, both with a hint

Performancerank II working

5 challenges, one profiler

Code reviewrank III fluent

6 diffs read, 1 blocked correctly

Designed progression, with example values.

Who it is for

Before the first job, in the first month of it, and after the agents moved in

The self-taught developer before the first job

Can build from a clean start and has never opened a codebase they did not write, with a deploy that fails and a report that names a symptom.

The team lead onboarding juniors

Wants the first month to be spent on the team's own codebase. Challenges built from the team's repositories, graded the same way, are the onboarding track.

The engineering lead whose team ships with agents

The code arrives faster than anyone reads it, and the lead can no longer say who could debug the billing path at three in the morning. The team track keeps that answer current.

The team track

Agents write the code. The team still has to understand it.

Short drills built from the team's own repositories and incident history, so the practice is on the systems each engineer is on call for, and it pays off at the next page rather than in a course certificate.

A team eighteen months into heavy agent use ships more code than it reads. Everything passes review, everything runs, and the skill that goes quietly is the one nobody measures: following a failure through code you did not write. It shows up as an incident that takes days, or a module that only one person could explain, and they have just handed in their notice.

Anthropic's own randomised study of junior engineers found that how they used the model mattered more than whether they did: the ones who asked it conceptual questions understood the code afterwards, and the ones who delegated the writing to it did not. That is a behaviour, and behaviours can be practised.

drill 01

Bug hunt

A realistic fault planted in a sandboxed copy of one of the team's own modules, by mutation rather than by hand, with the module's own tests as the judge.

trains: debugging in your code
drill 02

Explain-back

Walk through a flow in your own words. The grader has read the actual code and checks each claim against it, not against a model answer.

trains: knowing what you run
drill 03

Predict the break

A merged pull request, before its consequences. Say what fails in production and where, then see what actually did.

trains: blast radius
drill 04

Incident replay

A past postmortem re-run as a timed exercise, with the logs and the repository as they were that day and the ending withheld.

trains: the next page
drill 05

Review drill

An agent-written pull request with one real flaw in it, and every test green. Approve it or block it, and write down why.

trains: against the rubber stamp
cadence

Fifteen minutes, two or three times a week

Short enough to fit between tickets, and only worth doing if the team's lead has put the time in the week. Optional homework on a twelve-hour day does not happen.

the model

A tutor, never a solver

A model is on hand throughout, and it answers with questions: where would you look, what does that log line rule out. It has read the answer and will not write it.

the scores

Private to the engineer, pooled for the team

Each engineer sees their own results and nobody else does. The lead sees the team map below, which names modules, never people. A drill that feels like surveillance is a drill nobody takes honestly.

What the lead sees: modules, and how many people could take the page

checkout5 can explain it

explain-back and bug hunt, 14 drills this month

search indexer3 can explain it

one incident replay, median 41m to root cause

billing ledger1 can explain it

on the critical path; 80% of it merged from agents this year

auth gateway0 can explain it

no explain-back passed; review drills approved the planted flaw twice

Designed team map, with example values. A module at zero is the one to drill next, or to recover.

What it costs

Start free. Move up when the free challenges run out.

Four tiers. Individuals pay per seat, monthly; teams pay per seat for challenges built from their own repositories.

Free

free

Get your feet wet with real challenges.

  • Five challenges a month
  • Debug mode only
  • Community hints
  • Basic progress tracking

Mentor

per seat, monthly

A model reading over your shoulder, for faster growth.

  • Everything in Pro
  • Feedback from a mentor model
  • A personalised roadmap
  • Interview prep mode
  • Approach review and hints

Teams

per seat, teams

Onboard faster, and keep the codebase understood after the agents moved in.

  • Challenges built from your repositories
  • The team track: five drills from your own code and incidents
  • A team map by module, individual scores kept private
  • Onboarding tracks
  • SSO and admin controls
  • Dedicated support

Request access

Say who it would be for

Requests decide which stacks the first challenges are drawn from.

Free for the whole beta. The first twenty-five teams keep 50% off for twelve months after launch, locked in at signup.

Your address is used for this request and nothing else. One email when there is something to show.

Asked so far

Reasonable objections

There are interactive coding platforms already.
Exercism, Boot.dev and the interview platforms teach a language or an algorithm in a clean file, and do it well. Proddojo is the messy repository, the failing deploy and the logs: the situation a first week is made of, which none of them simulate.
A model writes the fix now.
It does, once the developer can describe the bug. Finding it in a system they did not write, from a report that names a symptom, is the part the model cannot do for them, and it is the part employers test.
Where do the bugs come from?
From open-source projects under permissive licences, with the report, the discussion and the fixing commit, credited on every challenge. The codebases are theirs; the environment, the grading and the hints are the work here.
Our code cannot leave the building.
Team-track drills are built in a sandboxed copy of the repository, and nothing from it is used for anyone else's challenges. For teams whose code may not leave their estate at all, the drill builder runs inside it, and only the scores come out.
Engineers will read this as surveillance.
They will if it is built that way, so it is not. Individual scores are visible to the engineer alone, there is no leaderboard on the team track, and the lead's view is by module: where understanding is thin, never who is thin on it.