# The Cut Test — Cross-Scorer Tracker

**Purpose:** One page of truth. Two independent scorers (Claude, GPT) evaluated the same evidence. Nothing goes on the public page unless it appears in Section 1. Updated July 25, 2026.

---

## 1. VERIFIED — both scorers agree, arithmetic/artifact checked

| Claim | Value | Verified by |
|---|---|---|
| Raw source duration | 18:31.048 (1111.0s); speech span 105.9–1097.3s | ffprobe + transcript span, both scorers |
| Reference edit | 2:17.731, 362 words (paid instrument) | ffprobe + paid measurement, both |
| Arm A (Claude) render | 2:28.235, 357 words; **+10.504s vs reference** | ffprobe + arithmetic, both |
| Arm B (TimeBolt auto) render | 2:43.4, 410 words; +25.8s vs reference | ffprobe + paid measurement (Claude scorer); GPT inventory |
| Take clusters | 15 clusters: **6 exact-match, 8 defensible alternatives, 1 clear error** | GPT source-mapping CSV; consistent with Claude scorer's 15-of-16 framing |
| The one clear take error (Arm A) | "You get 2, never 3" beat (GPT beat 11): kept version missing the "Fast, secure, watchable" lead | Both scorers, independently |
| Reference opening false starts | Deliberate stylistic beat, pre-registered in locked answer key, excluded from scoring | Answer key §3; both scorers applied it |
| Untranscribed-sentence mechanism | Sentence starts at source 08:47.94 (527.94s), inside Arm A KEEP 527.21–534.72; **preserved by timestamp accident, not semantic understanding** | GPT 813-word re-transcription + cuts.json; Claude scorer verified window fit (0.73s offset, 7.51s window vs ~7s sentence) |
| Input transcript defects | 741 words; 16 words >2s duration (max 19.996s); one full sentence missing. Stronger re-transcription: 813 words, zero >2s durations | Both scorers |
| Clipped word "IP" | Arm A render ends rant at "get my."; source "IP." was pre-registered poisoned timestamp idx 374 (4.6s bleed); KEEP boundary at 527.21 consistent with the clip | Paid measurement + answer key §5 + cutlist geometry |
| FFmpeg fairness | Arm A re-encoded all segments (H.264/AAC, 20ms fades, loudnorm, 4ch→stereo downmix); stream copy only at concat. **Keyframe snapping is not a confound** | GPT pipeline audit; consistent with run transcript |
| Arm A steelman confirmed | Own waveform stage: 10ms RMS frames, 200ms boundary search, per-channel gating, snapping, padding, fades. Methodology must NOT say "transcript timestamps only" | GPT audit; matches setup prompt requirements |
| Boundary distances | 50 paired boundaries; median 360ms from reference; 4/50 within 50ms, 9/50 within 100ms, 18/50 within 250ms. **Distance metric, not error count** — ears required | GPT boundary CSV (accepted; not independently recomputed) |
| Destructive errors, both arms | 0 confirmed at transcript level (paid instrument) | Claude scorer, corrected record |
| Measurement-layer fragility (Finding 4) | Free WhisperX: missed ~2min in raw; hallucinated "Thank you." over ~20s real speech (Arm A render); emitted "..." over intelligible speech (Arm B render). Paid instrument required for valid scoring | Claude scorer, cross-validated free vs paid |
| Reference used mixed-take constructions | Beats 11, 18, 30 stitch material from two takes (24+25, 42+43, 65+66). Arm A was barred from stitching by its own locked instruction | GPT CSV. **Must be disclosed on page**: composite editing was a human privilege by design |

| Audible boundary defects (ear-pass, final valid sessions) | Arm A: **20** (confirmed, incl. predicted "IP" clip); Arm B ladder: **B1 8 → B2 4 → B3 2**. Locked comparison remains A 20 vs B1 8 (2.5x); B2/B3 are addendum. Ablation: algorithm owns 8→2; aggressive threshold added 2→4 (the three transcript shaves); padding governs excision completeness. Scorer's padding-attribution prediction: wrong, retained | Two listeners, chopped-word scope, blind to render identity |
| Post-test tuning (v7.3.6) | Retake-detection code change: 3 duplicate-take classes eliminated, all content preserved, ear defects 8→4. Reported as addendum; locked-run numbers unmodified | Paid re-measure + ear re-count |

| Cleanup model revision (post-publication-draft) | Arm A revised 35–50 → **45–70 min modeled, floor ~45**. Canonical wording (GPT-refined, adopted): "modeled after mandatory source reconciliation and timeline reconstruction; visual repairs uncounted." Corrections applied: NLE export is NOT universally mandatory — fastest practical route requires either an NLE timeline or repeated pipeline revisions; watching the short output at 1x cannot reveal absences — defensible requirement is one complete source reconciliation (~18:31) or equivalent side-by-side audit. New scorecard row: "Repair context" (Arm A: flat render + cut manifest; Arm B: editable source timeline already present). Distance to publishable = discovery + repair + reconstruction. All cleanup figures remain labeled estimates | Founder floor argument + GPT wording audit, adopted by Claude scorer |

| Cost replication (July 26, API billing, claude-opus-5) | **VERIFIED:** metered recreation of the locked procedure — all 98 decisions, all 43 KEEPs reproduced; five boundaries varied 10–50ms ("decision-level reproducible, not byte-for-byte deterministic" — GPT wording adopted, corrects Claude scorer's "deterministic" overclaim). Ledger: 33/9,454/46,267/802,075 across 18 messages ≈ **$0.93 at list rates, matching Console**. **RETRACTED (Claude scorer):** the $9.16 one-word-repair figure, the 10x ratio, and the 20-defect extrapolation — published to the page draft before its session ledger existed; /cost output read "since start or last resume" (possibly cumulative); no second JSONL in project ledger; Console unchanged at time of audit. Removed from page; page carries the retraction openly. | Repair measurement — VALID (July 27, 00:10–00:17Z) | Second JSONL confirmed on disk (c9054ee9), separate from cost run (89909166). Protocol: defect reinstated (523.01, re-rendered), **confirmed clipped by ear pre-session**; headless Claude Code, same binary/key/protocol as $0.93 run; **confirmed intact by ear post-session** (boundary set 523.25 — tighter than reference 523.351, 79ms tail, word whole; next-segment audio immediately after is the splice working, not bleed). Ledger deduped: 62 in / 20,497 out / 51,648 cache-write / 1,479,721 cache-read, 33 msgs, ~7 min wall. **List-rate cost: $1.58 vs full edit $0.93 — repair = 1.7x the entire edit.** Labeled: recreation of historical repair; executing a known fix, not discovering one; no extrapolation to remaining 19 defects (batching possible). Claude scorer prediction ($0.10–$1.00, <8 min): time right, cost ceiling broken — retained. Console cross-check pending (~$2.51 expected); screenshot to vault when posted | GPT audit 28-metered-api-replication-audit.md, accepted with one factual supplement |

| Current-builds panel (July 26 regenerations, RESOLVES the B1-headline dispute) | Doug ear-counted the $0.93 replication render (edit_before_ip_fix.mp4, un-doctored full-edit regeneration) under the standardized pass: **Claude 17**. Vs TimeBolt B3 (v7.3.6, original settings): **2**. Panel added to page below locked table, clearly labeled "both re-run July 26, current builds": **17 vs 2, 8.5x**. Locked table (20 vs 8) unchanged above it. 20→17 drift corroborates decision-level-reproducible finding (5 boundary shifts + self-corrected IP). Resolution of three-day dispute: GPT's proposal to replace B1 with B3 in the primary scorecard REJECTED (asymmetric re-run = rigging accusation); Doug's underlying demand satisfied via symmetric regeneration of BOTH arms. Marketing may use 17-vs-2 / 8.5x for current-product claims; locked 20-vs-8 remains the pre-registered test result. **GPT convergence (final):** GPT retracted 20-vs-2 as run-mixing, adopted 17-vs-2 as canonical current-builds pair, endorsed no-new-test-before-shipping and Cut Test 2 (scripted + Zoom) post-launch. Structural difference retained by Claude scorer: locked table stays first, current-builds panel prominent beneath it (chronology carries authority; highlight carries attention). Panel counts labeled as founder's standardized listening counts per GPT's attribution requirement. All three parties aligned on all numbers as of July 27. |

## 2. CORRECTIONS LOG — kept in full, publishable

| # | Claim | Status |
|---|---|---|
| C6 | First ear-pass session (21–29 vs 8–11) invalidated: listeners counted different defect types (chopped words vs doubled takes). Re-run with single agreed category | **Superseded** by final session: 20 / 8 / 4 |
| C7 | (Claude scorer) "v7.3.6 eliminated 3 duplicate-take classes" | **Resolved by ablation (B3)**: algorithm removes only the opening duplicate, and incompletely at 0.01s padding (fragment residue). Triple and orphan never reduced — apparent reductions were settings-induced word shaves ("N"/"Leo"/"skip"), confirmed absent in B3 at original settings |
| C1 | (Claude scorer, free instrument) "Arm B dropped the rant + Descript passage; ≥3 destructive errors" | **Retracted** — instrument artifact; paid measurement shows all content present |
| C2 | (Claude scorer, free instrument) "Arm B retained ~58s dead air" | **Retracted** — zero gaps >3s on paid timeline |
| C3 | (Claude scorer) "Arm A kept the untranscribed sentence via the channel guard" → then "via phantom-duplicate KEEPs at 556–570s" | **Both wrong.** True mechanism per GPT + cutlist: accidental coverage by KEEP 527.21–534.72. Guard rescued only the 480–535s passages |
| C4 | (Claude scorer) "cuts.json contains two identical-text KEEP segments" | **Resolved** — truncated preview strings; GPT CSV shows 556.74–569.92 is one take (36) split by a micro-drop |
| C5 | (Claude scorer, early conversation) "Arm B is conservative, errors mostly slop" → revised against Showdown benchmark → **original prediction empirically correct on this file** | Restored |

## 3. OPEN DISCREPANCIES

None load-bearing. Cluster counts differ (GPT 15 vs answer key 16 scoreable) due to clustering granularity; conclusions identical. Terminology on the page: use GPT's 6/8/1 split — it is finer-grained and both scorers endorse it.

## 4. BLOCKING BEFORE PUBLICATION — all owner: Doug

1. ~~Ear-pass~~ **DONE** — final counts 20 / 8 / 4 in Section 1.
2. ~~Four numbers~~ **DONE** — reference <18 min; setup ~45 min wall / ~35 compute (reported as operator-dependent range); locked-run compute ≈⅒ setup, 8 min/render; cleanup rubric: Arm B 12–15 min in-timeline vs Arm A 35–50 min requiring source access. Human <1x runtime on every task.
3. **Model/version** of the Claude Code run.
4. **Locked-prompt artifact**: save as a file into the evidence set. MUST include the pre-lock additions verbatim — the four policy answers AND the channel-verification guard — or the methodology misrepresents the run.
5. **Arm B configuration statement**: silence settings (logged: −45dB / 0.5s / 0.75s / 0.01–0.15s pad) plus exactly which UMCheck removals were enabled.
6. **Deliver to GPT**: the locked answer key + scored results v1 (corrected) so both scorers cite one rule set.
7. **Evidence custody**: the 813-word stronger raw transcription and the paid measurement JSONs go into the public evidence set alongside the original 741-word input.

## 5. DECISION HYGIENE

The public page may state only Section 1 claims. Section 2 publishes as the corrections log — it is the credibility engine, not a liability. Any new claim by either scorer enters Section 1 only with artifact verification.

*TimeBolt, LLC — Keep It Real*
