Home I Let AI Agents Run a Nationwide CTF. The Hard Part Wasn't Solving the Challenges.
Post
Cancel

I Let AI Agents Run a Nationwide CTF. The Hard Part Wasn't Solving the Challenges.

I competed individually in Japan’s Ministry of Internal Affairs and Communications (MIC) nationwide CTF. The confirmed event rules permitted AI use. The official CTFd profile records 26th place and 5,701 points; the local tracker is a separate record.

Final CTFd profile crop showing 26th place and 5,701 points; per-challenge submission rows are excluded.
Final CTFd result: 26th place, 5,701 points. The published crop omits the per-challenge submission table.

Official CTFd result: 26th place · 5,701 points

Participation / rule: individual participation · AI use permitted

Local tracker: 110 tracked challenges · 107 recorded submit results · 101 locally recorded correct · 6 incorrect

Event)

The nationwide CTF was also an operations problem: collect, solve, verify, and submit many challenges within a fixed event window. This is a pipeline retrospective, not a challenge-answer write-up. I distinguish the official result from local tracker arithmetic and observed event traces from what the retained evidence cannot establish.

Result)

The preserved tracker contains 110 challenges and 107 recorded submit results. Of those results, 101 are marked correct and 6 incorrect; three tracker rows have no recorded submit result. A result-row count is not a count of automated solves.

The 101 locally recorded correct rows sum to a nominal 5,900 tracker points. That is not the official score. The 199-point difference from the CTFd profile remains unexplained by the retained evidence, so I keep the figures separate.

Category roll-up of 110 tracker rows: 107 recorded results, 101 correct, 6 incorrect, and 3 with no recorded result, plus category and state counts.
Local summary from the preserved tracker, not the official scoreboard. Recorded-result coverage is not an automation rate.

Approach)

The local evidence contains 110 challenge records and 21 attachments across 20 challenge IDs. Session logs show both local advisor/model activity and remote API use. They also contain directions to work in parallel and a separate submission worker.

These records establish artifacts and some event activity, but not a complete end-to-end implementation. They do not map each challenge to a worker or model, identify human interventions, or preserve retry history.

The author-provided event capture records an intermediate CTFd profile at 29th place and 4,401 points. The OMP pane and per-challenge history from that capture are withheld because they expose challenge-specific operations. An earlier capture at about 13:53 JST showed 302nd place and 1 point; that was an early state, not the final result in Figure 1.

Intermediate CTFd profile crop showing 29th place and 4,401 points; the OMP pane and challenge history are omitted.
Intermediate CTFd state during the event: 29th place, 4,401 points. This is not the final result.

The supplementary timeline below is reconstructed from preserved control-plane events and intervention records. It extends the chronology but is derived evidence, not an OMP screenshot or a measure of active solve time.

Supplementary derived UTC event timeline; overlapping task lifetimes do not represent active solve time.
Supplementary, derived event timeline from preserved control-plane events and intervention records. Overlapping task lifetimes are not active solve time.

Problem)

The tracker does not describe the whole pipeline. Owner fields are blank, and challenge start/end times do not provide a complete event history. HTTP transport success is separate from the challenge application’s semantic result.

If queue state is copied between workers and submission results arrive late, work can be duplicated or continue after it should have stopped. The retained evidence does not establish the cause of every operational symptom.

Bottleneck)

The clearest operational bottlenecks in the retained evidence are in the control plane.

  • Ownership and state: without a per-challenge owner and lease, duplicate work and responsibility boundaries cannot be measured.
  • Artifact intake: the tracker manifest and downloaded files cannot be assumed to match. Intake needs hashes and provenance.
  • Verification: check result format and meaning independently before submission.
  • Observability: challenge-keyed event records are needed to separate model performance from queue and control-plane delay.
  • Close control: without an authoritative cutoff and official result export, late work is hard to stop safely and the tracker is hard to reconcile.

Local vs API)

Without per-challenge route data, I cannot estimate how many API calls local compute avoided or how many challenges required human intervention. An exact automation percentage is also unavailable.

Architecture v2)

For the next run, I would use one canonical queue with an active lease keyed by challenge ID. Intake would validate file types and hashes, followed by deterministic local analysis. Only then would work route to a health-checked local model or a remote API with an explicit budget. Candidate material would remain in a private evidence store. An independent verifier would gate a single submission broker, and an official cutoff barrier would close submissions before score reconciliation.

This is a proposal, not a description of the 2026 implementation. The left panel is a V1 reconstruction bounded by retained event evidence, not a complete implementation specification. The right panel is a proposed V2, not a comparison of completed implementations.

The left panel reconstructs V1 artifacts and unknown links from event evidence; the right panel shows the proposed V2 control-plane architecture.
Evidence-bound V1 reconstruction (left); proposed V2 control plane (right).

Retrospective)

Next time, I would make per-challenge execution and submission state observable before adding workers.

Evidence / Notes)

  • Figure 1 is the final CTFd evidence: 26th place, 5,701 points.
  • The local tracker records 110 challenges, 107 results (101 correct, 6 incorrect, 3 unrecorded), and 5,900 nominal points. This is separate from CTFd; the 199-point difference remains unexplained.
  • Figure 4 is a CTFd-only crop of an author-provided event capture at 29th place / 4,401 points; the associated OMP pane and per-challenge history are withheld. The earlier 302nd-place / 1-point state at about 13:53 JST was also an early snapshot, not final.
  • The supplementary timeline is derived from retained events. Figure 5 V1 is separately reconstructed from retained evidence; V2 is proposed.
  • Per-challenge model routes and human-intervention counts remain incomplete. No precise automation percentage is claimed.
This post is licensed under CC BY 4.0 by the author.

Buy me a coffee