# The Cut Test — Answer Key (Pre-Registered)

**Locked:** July 25, 2026 — before either automated arm has run.
**Author:** Doug Wulff, TimeBolt, LLC
**Purpose:** This document inventories every editorial decision in the raw recording, the reference edit's resolution of each, and the predictions registered before any automated output existed. It is published unmodified alongside the test results.

---

## 1. Source Material

| Item | Value |
|---|---|
| Raw recording speech span | 105.9s – 1097.3s (~16.5 min) |
| Raw transcript words | 741 |
| Effective speech density | ~45 words/min (dead-air heavy by design of the recording session, not by staging) |
| Dead-air gaps > 4s | 40 |
| Reference edit runtime | 2:11 (131.5s) |
| Reference edit words | 347 (47% of raw words, 13% of raw span) |
| Transcript format | Word-level JSON, start/end/word schema |
| Fillers preserved in transcript | Yes — 7 um/uh instances, false starts verbatim |

**Fidelity verdict:** The transcript is disfluency-honest. Arm A receives a fair input. Engine that produced this JSON: [CONFIRM: TimeBolt Transcribe or WhisperX] — to be stated in methodology.

---

## 2. PREFLIGHT FINDING #1 — Reference contains speech absent from the raw transcript

The reference edit includes the sentence:

> "TimeBolt solves the cleanup trap locally by cutting the signal itself, skips silence like it never even happened."

The words "solves," "skips," and "cutting the signal" appear **nowhere** in the 741-word raw JSON. Possible explanations, to be resolved before Arm A runs:

- (a) The raw transcription missed a spoken segment (candidate regions: the silent 0–105.9s lead-in; the 772–897s gap; 902–926s; 946–1001s)
- (b) The reference incorporated material from a different source take/file
- (c) Transcription variance between the raw pass and the reference-render pass

**Resolution:** [FILL IN before Arm A runs]

**Why it matters:** If (a), the raw JSON is incomplete and Arm A structurally cannot keep speech it was never shown — a transcript-pipeline failure mode worth publishing in its own right. If (b), the sentence is excluded from all scoring. If (c), the methodology notes the variance.

---

## 3. Take Inventory — 19 Decision Points

"Correct automated behavior" = what any competent editor would do. Take *identity* (which of several clean takes survives) is a **CHOICE**, never an error, unless the surviving take is objectively flubbed. Items marked **STYLISTIC** contain deliberate reference-edit beats no automation is expected to reproduce; those beats are excluded from error scoring.

| # | Content | Raw location (word idx / time) | Takes | Correct automated behavior | Reference resolution | Scoring notes |
|---|---|---|---|---|---|---|
| 1 | "How long is it?" | 0–31 / 105.9–165.1s | 5 | Keep exactly 1 | Kept 1 | Take 3 (idx 8–11) contains poisoned timestamps |
| 2 | "It's the first question everyone asks" | 12–37 / 133.8–172.4s | 3 | Keep exactly 1 | Kept 1 | Interleaved with #1 |
| 3 | "Hey team, so, um, this quarter, we…" cascade | 38–106 / 221.7–277.2s | 6 false starts, **none completes** | Remove entire cascade | Kept one fragment ("I had it a second ago. Let me start over," idx 86–95) as a deliberate beat | **STYLISTIC** — the kept fragment is excluded; removing everything scores as correct |
| 4 | "You don't want a 10-minute recording / just the six that matter" | 107–138 / 318.2–340.3s | 2 complete + 1 partial | Keep 1 complete pair | Kept 1 | Partial (idx 120–125) is junk |
| 5 | "But how do you make internal video watchable…" | 139–153 / 345.0–350.2s | 1 | Keep | Kept | No decision |
| 6 | "Fast, secure, watchable. You get two, never three" | 154–178 / 359.4–396.6s | 4 attempts (1 lacks "never three," 1 partial) | Keep 1 complete incl. "never three" | Kept 1 complete | Takes 3–4 contain poisoned timestamps |
| 7 | "Talk to a screen… training never/didn't happen" | 179–216 / 402.7–425.2s | 2 | Keep 1 | Kept fragments of **both** as echo | **STYLISTIC** — echo excluded; 1 take = correct |
| 8 | "Edit by hand… burn an employee" | 217–240 / 433.0–444.5s | 2 near-identical | Keep 1 | Kept 1 | Pure choice |
| 9 | "Someone reaches for cloud AI… what it cut" | 241–272 / 449.4–466.1s | 2 | Keep 1 | Kept 1 | Pure choice |
| 10 | Problem block (Loom/Descript…servers) | 273–335 / 476.7–505.9s | 1 pass + embedded junk | Keep, strip junk (see §4) | Kept, junk stripped | Junk: "Yeah," / "So, like," / "Yes. Oh, shit." / abandoned "You've heard the pitch without…" |
| 11 | Enterprise rant ("chill wax… my IP") | 336–374 / 513.9–534.6s | 1 | Keep, strip "Okay." and "uh," | Kept, stripped | Sounds unscripted; deleting it entirely = **destructive**. Contains 2 poisoned timestamps. Note: "chill wax" is a transcription artifact (likely "chillax") |
| 12 | Three-layer explanation | 375–437 / 550.0–599.6s | Attempt 1 (line only), attempt 2 (complete), false start, attempt 3 (complete) | Keep 1 complete | Kept 1 | Poisoned "difference." (5.7s) sits at attempt 2's end |
| 13 | "First, the cleanup trap" + full re-take of #12 | 468–510 / 639.9–669.3s | 1 line + 1 complete re-take | Either #12 or #13 block survives, not both | Reference used cleanup-trap framing (see §2 finding) | Keeping both blocks = slop |
| 14 | "Timebolt is a lightweight app…control" | 438–467, 511–562 / 606.4–703.8s | 3 complete + 2 false starts + 1 orphan "Timebolt." | Keep 1 complete | Kept 1 | Most repeated block in the file; poisoned "Mac…" (3.8s) |
| 15 | "One engine, every way your company records" | 563–582 / 710.7–728.7s | 2 (take 1 poisoned: "records." 6.9s) | Keep 1 | Kept 1 | Take 2 carries the continuation (async training…) |
| 16 | "11,700 paying customers… Vault" | 598–615 / 754.4–772.0s | 2 (take 1 ends in abandoned "Vault…" 6.2s poisoned) | Keep take 2 | Kept take 2 | Take 1 is objectively incomplete — keeping it = error, not choice |
| 17 | "And this video, captured and edited entirely inside Timebolt" | 616–624 / 896.8–901.8s | 1 | Keep | Kept | Follows the 124.8s silent region |
| 18 | "The product made the pitch" | 625–636 / 925.9–973.6s | 1 false start + 2 complete | Keep 1 | Kept 1 | False start contains the worst poisoned word in the file ("product," 20.0s) |
| 19 | Closing block + "Keep it real" | 637–740 / 1000.9–1097.3s | "Finally, all three" ×2, "Pick three" ×1; closing ×4; "keep it real" ×5 | Keep 1 tagline variant, 1 closing block, 1 "Keep it real" | Kept "Pick 3." + final closing + 1 KIR | Choosing "Finally, all three" over "Pick three" = defensible CHOICE. Take 3 of closing contains poisoned "leaving" (13.1s) |

**Scorecard denominator:** 16 scoreable take-selection decisions (19 minus #5 and #17, no decision; #3 scored as remove-all).

---

## 4. Junk Inventory (must-remove items outside take selection)

| Item | Word idx | Time | Class if retained |
|---|---|---|---|
| "um…" / "um," ×6 | 41, 45, 56, 67, 82, 99 | 222.9–273.2s | Slop (inside cascade #3 — removed with it) |
| "uh," | 357 | 519.9s | Slop |
| "Yeah," | 279 | 479.8s | Slop |
| "So, like," | 291–292 | 484.7s | Slop |
| "Yes. Oh, shit." | 311–313 | 490.3–491.0s | Slop (profanity — mandatory removal for publishable) |
| "Okay." | 336 | 513.9s | Slop |
| Abandoned "You've heard the pitch without…" | 318–322 | 498.3–500.5s | Slop |
| Orphan "Timebolt." | 438 | 606.4s | Slop |
| Partial repeats (idx 120–125, 160–162, 419–422, 535–538, 625–626, 675–677) | — | — | Slop |

---

## 5. PREFLIGHT FINDING #2 — 16 Poisoned Timestamps

These words carry end-timestamps that absorbed trailing silence (alignment bleed). Any pipeline cutting on raw `end` values will embed dead air inside "kept" words. Registered before Arm A ran.

| Word idx | Word | Start | End | Bleed |
|---|---|---|---|---|
| 9 | long | 116.14 | 123.20 | 7.1s |
| 11 | it? | 123.48 | 129.15 | 5.7s |
| 41 | um… | 222.95 | 226.31 | 3.4s |
| 48 | we… | 229.98 | 232.74 | 2.8s |
| 85 | we… | 261.86 | 264.28 | 2.4s |
| 164 | secure, | 377.32 | 381.35 | 4.0s |
| 171 | fast, | 385.62 | 391.43 | 5.8s |
| 370 | they're | 522.17 | 527.86 | 5.7s |
| 374 | IP. | 529.96 | 534.55 | 4.6s |
| 390 | processed. | 560.87 | 563.97 | 3.1s |
| 418 | difference. | 582.44 | 588.09 | 5.6s |
| 443 | Mac… | 611.07 | 614.86 | 3.8s |
| 569 | records. | 712.99 | 719.84 | 6.9s |
| 603 | Vault… | 758.54 | 764.79 | 6.2s |
| 626 | product | 926.13 | 946.13 | **20.0s** |
| 714 | leaving | 1064.16 | 1077.22 | 13.1s |

Note the distribution: poisoned words cluster at take boundaries and abandoned fragments — precisely where cut decisions concentrate.

---

## 6. Structural Regions

| Region | Duration | Content |
|---|---|---|
| 0 – 105.9s | 1:46 | No transcribed speech (lead-in) |
| 772.0 – 896.8s | 2:05 | No transcribed speech (demo/product segment) |
| 901.8 – 925.9s | 24s | No transcribed speech |
| 946.1 – 1000.9s | ~55s (interrupted) | Sparse fragments + silence |

How each arm handles the two large no-speech regions is a scored line item. These are also candidate locations for the §2 missing sentence.

---

## 7. Scoring Rules (fixed)

Three buckets, per arm, reported separately. **No composite score, no numeric weighting — repair time is the weight.**

- **RECOVERABLE SLOP** — should have been removed, remains. Repair: find, delete.
- **DESTRUCTIVE ERROR** — should have stayed, was removed. Repair: return to source, restore, rebuild splice, verify.
- **BOUNDARY ERROR** — right decision, bad splice (onset shaved, breath cut, gap, click). Repair: nudge, replay.

Additional fixed rules:
1. Take identity among clean takes = CHOICE, not error.
2. Stylistic reference beats (#3 fragment, #7 echo) are excluded from error scoring in both directions.
3. "Publishable" means a publishable pitch video of this recording's content — not a replica of the reference timeline.
4. Repair tooling is inside the measurement: Arm B repairs occur in the TimeBolt timeline; Arm A repairs occur via re-prompting or manual NLE work.
5. Both arms scored with the same ink. Prior published TimeBolt benchmark (90.6% F1 bad-take removal, scripted datasets) predicts Arm B will not be perfect.
6. If either arm performs well anywhere, it is stated plainly. Results publish regardless of outcome.

---

## 8. Registered Predictions (before any automated output existed)

1. **Poisoned timestamps** will cause Arm A to embed dead air inside kept segments at one or more of the 16 flagged words, worst at idx 626 ("product," 20.0s), unless its pipeline snaps boundaries to the waveform.
2. **Take selection:** Arm A will resolve some of the 16 scoreable decisions incorrectly — most likely failure sites are #13 (block-level duplicate: keeping both three-layer explanations) and #16 (keeping the objectively incomplete take 1).
3. **Boundary quality:** Arm A cuts made directly on word timestamps will shave word onsets; Arm B (L3 correction) will not, per prior benchmark.
4. **Arm B** will leave redundancy consistent with its published 85.5% cleanliness — expected form: at least one duplicate take pair surviving among items #8, #9, #14, #19.
5. **The silent regions** will be handled correctly by both arms (trivially detectable), but the sparse-fragment region (946–1001s) may confuse Arm A's notion of segment continuity.

Predictions that fail are reported as failed.

---

*This document was locked before either automated arm produced output. The reference edit was completed and locked before its author viewed any automated output. Active editing time for the reference edit: [FILL IN] minutes.*

*TimeBolt, LLC — Keep It Real*
