# The Cut Test — Scored Results (v1)

**Scored:** July 25, 2026, against the pre-registered answer key (locked earlier the same day, before either automated arm ran).
**Measurement instrument:** Paid-tier WhisperX-class transcription, identical engine and config across all three renders. (Free local WhisperX large-v3 was the original instrument and was replaced mid-scoring after cross-validation revealed transcription failures — see Finding 4.)

---

## The three files

| | Reference edit | Arm A — Claude autonomous | Arm B — TimeBolt automated |
|---|---|---|---|
| Runtime | 2:17.7 | 2:28.2 (+10.5s) | 2:43.5 (+25.8s) |
| Measured words | 362 | 357 | 410 |
| Produced by | Doug Wulff, hand-edited in TimeBolt from the raw audio | Claude Code + WhisperX word JSON + FFmpeg, locked prompt, zero human intervention after lock | TimeBolt silence detection (−45dB, >0.5s, ignore <0.75s, pad 0.01/0.15s) + UMCheck, default settings, no timeline touch |

## Headline scorecard

| Metric | Arm A (Claude) | Arm B (TimeBolt auto) |
|---|---|---|
| Take-selection decisions correct (of 16) | **15** | **~12** |
| Recoverable slop retained | **2 items** ("You've heard the pitch of that" fragment; "You get it" fragment) | **~7 items** ("How long is it?" ×2; "You get 2 never 3" section ×3 versions; both training takes; "Finally all three" AND "Pick 3"; waveform false-start fragments; "2019," orphan; "video leaving your control" orphan) |
| Destructive errors (content lost) | **0 confirmed** | **0 confirmed** |
| Audible boundary defects (ear-pass, chopped-word scope) | **20** — including "IP," clipped at a pre-registered poisoned timestamp (answer key §5, idx 374) | **8** (original locked run); **4** after post-test retake/boundary update, v7.3.6 — see addendum |
| Dead air >3s retained | 0 | 0 |
| Kept the untranscribed sentence | Yes — **by accident**: misaligned transcript timestamps caused a KEEP window to capture unclaimed audio (verified in cuts.json; independently identified by a second scorer, GPT) | Yes — by architecture (L1 keeps audio regardless of transcript) |
| Publishable without cleanup | No | No |

**Both arms reached "no." The reference reached "yes." Distance to publishable is the story, and the two arms fail in opposite directions:** Arm A is aggressive and mostly right — near-reference runtime, 15/16 judgment calls, two fragments of slop, one clipped word at exactly the timestamp defect we flagged in advance. Arm B is conservative and safe — zero content lost, everything real preserved including speech no transcript contains, at the cost of roughly a dozen duplicate takes and fragments, every one of which is a two-keystroke deletion in the timeline that produced it.

---

## The four findings

### Finding 1 — The transcript is missing real speech, and only architectures that ignore the transcript survived it
WhisperX transcribed 741 words from the raw recording and missed at least one full spoken sentence ("TimeBolt solves the cleanup trap locally by cutting the signal itself, skips silence like it never even happened") plus surrounding speech in a ~2-minute region it returned as empty. The human editor kept that content (working from audio). TimeBolt's automated pass kept it (L1 operates on the waveform; audio present = audio retained). Claude's pipeline kept it **by accident**: the stronger transcription places the sentence's start at source time 08:47.94 (527.94s), inside Arm A's KEEP segment at 527.21–534.72 — a window labeled with rant-tail text whose bleeding word-timestamps spanned that audio. The sentence survived incidentally inside a keep the pipeline made for other words. (Mechanism identified by the second scorer, GPT, from its 813-word re-transcription of the raw file; verified against cuts.json. The first scorer's initial attribution to the channel-verification guard, and a subsequent phantom-duplicate hypothesis, were both wrong and are retained in the corrections log.) This is preservation by timestamp accident, not semantic understanding — the pipeline could just as easily have lost it. The channel-verification guard, separately, is what rescued the rant and Descript passages at 480–535s.

### Finding 2 — The near-catastrophe: confident, well-engineered, wrong
During setup, the Claude pipeline determined that two windows totaling 33 seconds (the Loom/Descript passage and the enterprise rant — the strongest unscripted material in the recording) were "digitally silent" and that the transcript's 73 words there were fabricated. Its evidence chain was locally sound at every step: channel probing, amplitude checks on its selected voice channels, zero-peak confirmation across two windows, printed "definitive." It built a guard and deleted the content in its setup render. The speech was real — it lived on the audio channels the pipeline had ruled out from sampled checks. A human noticed the misdiagnosis pre-lock (registered in the answer key addendum before the official run); one generic engineering requirement — "verify silence on every channel before classifying a span as fabricated" — was added; the locked run then caught its own error and kept the content, correctly dropping "Oh, shit" as a false start and "Okay"/"uh" as filler within it.
**The exhibit:** an autonomous pipeline manufactured rigorous-looking evidence for destroying 10% of the content, and was saved by one externally imposed verification rule that exists only because a human was watching.

### Finding 3 — The poisoned timestamp claimed its predicted victim
Sixteen words in the transcript carried end-timestamps that absorbed trailing silence (up to 20.0s), registered pre-run as predicted failure sites. Arm A's cut at flagged word idx 374 ("IP.", 4.6s bleed) clipped the final word of the rant: the render ends the sentence at "they're going to get my." The waveform knew where the word ended; the timestamp did not; the pipeline trusted the timestamp.

### Finding 4 — The measurement layer itself failed three ways
Free local WhisperX large-v3, used as the initial scoring instrument, (a) missed ~2 minutes of real speech in the raw file, (b) hallucinated "Thank you." over ~20 seconds of real, intelligible content in Arm A's render (One engine / SharePoint / 11,700 section — all present, confirmed by paid re-measurement), and (c) emitted "..." over intelligible speech in Arm B's render. Preliminary scoring based on the free instrument wrongly charged Arm B with three destructive errors and ~58s of dead air; both charges were retracted when the paid instrument was applied uniformly. The transcript layer the AI-editing ecosystem treats as ground truth is fragile even as a *measuring* device.

---

## Registered predictions — outcomes

| # | Prediction (locked pre-run) | Outcome |
|---|---|---|
| 1 | Poisoned timestamps cause Arm A to embed dead air or misplace cuts, worst at flagged words, unless it snaps to waveform | **Partially confirmed** — Arm A implemented waveform snapping (no dead air embedded), but clipped "IP" at flagged idx 374 |
| 2 | Arm A resolves some of 16 decisions wrong; likely #13 (duplicate block) and #16 (incomplete take) | **Mostly wrong** — Arm A got #13 and #16 right; its one take error was #6 (kept version missing the "Fast, secure, watchable" lead) |
| 3 | Timestamp-cut boundaries shave onsets; Arm B (L3) does not | **Half confirmed** — Arm A: 20 audible chopped words incl. the predicted "IP" clip; but Arm B was not near-zero (8 audible defects). Prediction overestimated L3 on this footage |
| 4 | Arm B leaves redundancy consistent with 85.5% cleanliness — at least one duplicate pair among #8, #9, #14, #19 | **Confirmed and exceeded** — duplicates at #1, #6, #7, #19 |
| 5 | Both arms handle silent regions correctly; sparse region may confuse Arm A | Silent regions handled; no confusion observed |
| 6 (addendum) | Arm A's final output will be missing the 514–535s content as a destructive error | **Wrong in the best way** — the channel guard (added post-registration, pre-lock) reversed it; disclosed as such |

Scorer's corrections, kept in the record: the scorer (Claude, this conversation) initially declared Arm A's channel finding a misdiagnosis (right), then declared Arm B had dropped the rant and retained ~58s of dead air (wrong — instrument artifact, retracted), and revised its Arm B slop model twice. All three reversals trace to instrument reliability, which is itself Finding 4. A fourth correction: the scorer initially attributed Arm A's retention of the untranscribed sentence to the channel-verification guard; a second scorer (GPT, run in parallel by the author) identified the true mechanism as accidental retention via misaligned timestamps, verified against cuts.json. Two-scorer disagreement resolved by artifact evidence is retained here as part of the method.

---

## Addendum — Post-test tuning (not part of the locked comparison)

The locked test found TimeBolt's own automated pass leaving duplicate takes (~7 slop items) and 8 audible boundary defects. Within 24 hours, a retake-detection and boundary update shipped to all users as **v7.3.6** (automatic; no user action). Re-measured on the same raw file, same paid instrument, same ear protocol:

- **Ablation (B3: new algorithm, original settings)** isolated the two changes. Algorithm effect: the opening duplicate is marked and mostly excised — but at 0.01s left padding the excision is incomplete, leaving a shaved fragment ("How is it?"); at 0.1s padding (B2) the removal is clean. All other duplicate classes (the "never 3" triple, both training takes, both endings, the orphan fragment) are unchanged by the algorithm — earlier claims of their reduction are retracted (corrections C7). Settings effect: the 0.3s min-silence threshold in B2 introduced at least three word shaves absent in both B1 and B3 ("never"→"N", "video"→"Leo", "skips"→"skip"). Threshold aggressiveness trades pacing for word damage; padding governs excision completeness.
- All major substantive sections preserved in every run; zero destructive at section level; dead air 0
- Audible boundary defects, full-render ear count: **B1 8 → B2 4 → B3 2.** Ablation attribution: the algorithm owns the boundary improvement (8→2 at identical settings); the aggressive 0.3s threshold *added* defects (2→4, matching the three transcript-visible shaves); padding governs excision completeness only (B3's residual "How is it?" fragment vs B2's clean removal). Scorer's registered prediction that padding drove the improvement was wrong and is retained as such.

The original locked-run numbers stand above, unmodified. Silence-detection slider changes (min silence 0.5→0.3s, left pad 0.01→0.1s) do not affect take selection; the duplicate reduction came from the retake-detection code change.

**Ear-pass method:** two listeners, second listener blind to which render was which; single defect category agreed aloud before playback ("words chopped or splices that pop — anything heard twice does not count, the transcript already counted those"); one play-through per render, tally with rough timestamps. An earlier listening session was invalidated because the two listeners counted different defect types (corrections log C6).

---

## Pending before publication

1. **Ear-pass, both renders** — splice quality at every cut point (clipped onsets, clicks, cadence breaks). Only remaining scored dimension. Includes confirming the "IP" clip audibly and Arm B's L3 boundary quality.
2. **Arm B configuration statement** — confirm which UMCheck removals were enabled (filler removal evidently ran; bad-take removal behavior on this file must be named precisely).
3. ~~Reference edit active minutes~~ **DONE**: <18 min (founder-reported actual — faster than the 18.5-min recording; the one measured human number). Arm A setup: not isolated; the ~45 min wall / ~35 min compute session reconstruction cannot substitute for active setup. Locked-run final encode ~8 min. Cleanup: Arm B <18 min (founder estimate, in-timeline); Arm A modeled 35–50 min requiring source access — modeled, not performed, and labeled as such everywhere. Token-cost self-report stays in the evidence record but off the page (no billing record).
4. Reference-edit bias disclosure, consistency note vs. the published Showdown benchmarks, and "reference edit" terminology per the locked answer key — carried into the page unchanged.

---

*Scored against answer key locked 2026-07-25 prior to any automated output. Both automated outputs published untouched. TimeBolt, LLC — Keep It Real*
