# v0.2 review and iteration record

Recorded 2026-10-01. All eight role reviews are **Codex simulations**. They combine role-based inspection with executable protocol probes. Observed human participants: **0**; observed human sessions: **0**. Invitations sent: **0**. There are no actual interview quotations or independent usability measurements in this record.

## Invitations and observation protocol

Simulation scenario: a small team wants to verify a first MCP task and decide which onboarding material to revise. The simulated invitation asks the role to inspect the documented setup, complete its assigned task, inspect acceptance, and explain the next maintenance action. Role probes are executed by Codex; their results measure the available mechanism.

For a future human study, invite a voluntary maintainer for a workflow interview or a developer for a first-task observation. Explain the task, collected data, withdrawal option and publication choice; request separate permission for recording and quotations. Store contact details and raw notes privately. Count a participant once across roles; count completed sessions separately. Record assistance, prior experience, version, task and outcome. An invitation or automated browser session contributes zero observed participants.

## Eight sequential reviews

| ID | Simulated role | Task and starting point | Executed check | Obstacle / review concern | Follow-up |
|---|---|---|---|---|---|
| M1 | AI API maintainer | Built-in task-01; B guide | Protocol execution, feedback, new guide, re-execution, revision handoff | Model and protocol denominators need separation | Separate mode totals and filters |
| M2 | MCP tool maintainer | Registered Filesystem container; docs-evidence-01 | Same-task execution and independent re-verification, changed guide and handoff | Registration and container requirements need a discoverable path | Link the external task guide from the tutorial |
| M3 | Documentation owner | task-07; existing B guide | Immutable material lineage, same-condition revision and handoff | A saved material comparison needs precise terminology | Describe matrix comparison and feedback revision separately |
| D1 | New Python/MCP developer | task-01; documented protocol workflow | Successful protocol task and independent acceptance | Default combined evidence can obscure first value | Default public view to model evidence with separate protocol selection |
| D2 | Experienced MCP developer | task-05; fixed task contract | Acceptance and allowlisted report creation | Saved pairs need condition and responder checks | Pair only matching execution conditions and known responders |
| D3 | Windows developer | task-09; isolated Python 3.11 | UTF-8 artifacts and re-verification | Historical evidence needs a stable entry | Add version selection and archived download |
| D4 | Container environment developer | docs-evidence-01; pinned offline image | Filesystem task and independent output rules | Deployment instructions should match current export validation | Update versioned build guidance |
| D5 | English documentation developer | task-11; English instructions | Correct invalid-input rejection and re-verification | Report-loading failure needs a retry path | Add reload control and clear loading/error states |

The private executable record contains each run ID, outcome, source label and verified revision ID. The three maintainers each perform one feedback-to-revision protocol loop; the five developers perform one task probe each. Normal output and correct input rejection follow independent task rules. Task probes use CPU; no new model inference is needed.

The concern column records simulated review judgments. It contains no claimed human observation or invented timing. Before/after evidence comes from source inspection, public-site checks, automated browser acceptance and the protocol probes.

## Decisions and retest

| Issue | User goal | Priority | Change | Acceptance evidence |
|---|---|---|---|---|
| R1 Mixed execution totals | Interpret model behavior separately from service correctness | High | Separate mode/cohort/split controls and totals | Browser checks both denominators |
| R2 Loose saved-case pairing | Trust the illustrated comparison | High | Matching condition fingerprint and responder; unknown responder excluded from walkthrough | Browser injects mismatched and unknown candidates |
| R3 Historical report replacement | Revisit earlier evidence | Medium | Archive v0.1, show v0.2 by default | Both downloads validate; browser switches versions |
| R4 Export numeric validation | Recompute finite cost bounds | High | Reject NaN/infinity and unexpected public fields | Public-release tests |
| R5 Setup and loading guidance | Find the next executable step | Medium | External-task link, reload state and updated deployment guide | CLI/container and bilingual browser checks |

Protocol execution validates plumbing and output contracts. It does not measure the effect of guide wording on a person. The frozen real-model factorial remains unchanged; its separate report provides material-effect evidence.

## Closed loops and next evidence

Engineering: task contract → execution → independent acceptance → failure evidence → material revision → re-verification → release checks.

Product iteration: role-based review → explicit concern → prioritized change → regression acceptance. This is a simulated/internal product loop.

External validation remains open: voluntary recruitment → actual workflow observation → supported issue → revision → observed follow-up. Keep this stage pending until records exist. A future resume may cite implemented mechanisms, model experiments and engineering results; observed adoption and time improvements require human evidence.

## Hosted integration follow-up

The first Linux CI attempt built the container successfully but failed output acceptance because its nonroot user could not write the host-owned output bind. The correction keeps the host run ancestor private and permits writes only through its dedicated sticky output directory; container network, root filesystem and input restrictions remain in force. The release requires a successful Linux protocol/isolation retest. Hosted browser checks fetch reports within the browser to use the same network path as the site.
