Weekly Review · Totem Terminal
2026 · Week 31 · Generated 2026-08-07

The week the Black Flag pitch was rejected twice before the process ran in the right order.

Sixteen sessions, one on every day of the week, built the same investor deck three times and learned that a report and a pitch are opposing forms. Alongside it, three separate threads turned up the same shape of problem: a gate that had been checking nothing.

The Numbers

Sessions
16
Hours
79h
Days Active
7/7
Commits
67
Files in Git
2,389
Deploy Targets
11
Pitch Versions
3
Docs Restyled
1,670

The first unbroken week in the record: every day from 2026-07-27 to 2026-08-02 carries both a daily note and at least one session record. Three sessions ran past eight hours, and the two longest crossed a date boundary overnight. Of the 2,389 unique paths git saw, 1,670 are the style-router bulk write on 2026-08-01.

Wins by Project

Black Flag investor pitch Jul 28 – Aug 1 · 5 sessions · 1,805 min

v2 built and rejected. The deck was built twice on 2026-07-28 and got the register wrong both times. The intake had been there the whole time; the rebuild that worked started from it.

The full ship sequence ran on v2 on 2026-07-29: grammar gate to 0 BLOCK and 0 WARN, Playwright clean at 1280, 768 and 390, a six-auditor megaeval producing 179 findings with 103 applied through a triaged ledger, adversarial verify to zero, deployed gated to the blackflag-narrative alias, and curl-verified 401 unauthenticated on the hash URL, the alias, and root.

The 2026-07-29 call rejected v2 outright. Before that call: a third megaeval round, the 15:00 working-session page, the redline Google Doc for KG, ShurIQ placed into its own competitive stack ranking at 39.9, and the deck loaded into and then removed from Report Studio.

v3 rebuilt from scratch that evening with the intelligence process run before the build rather than after it.

  • 85 traced directives, a 243-statement typed graph, three structural gaps
  • Eight scored negative-space claims; the competitive set rebuilt as categories
  • Shipped gated and verified, with Report Studio and a redline Doc alongside
  • Two animated Slack agent sequences built in HyperFrames and deployed as an ungated review page, after a byte-range defect that made the video impossible to scrub was found and fixed
  • KG's copy notes applied on 2026-07-31: judgment stanza, CarbonArc value, named customers, positioning sentence; first two team bios shipped; Gap Radar viewport dropped so the Brand Power Index leads with the funded projection
A confidentiality leak closed on 2026-08-01. Two Black Flag pages were serving publicly and were regated. Cloudflare Pages resolves functions/ from the working directory, so deploying from the parent ships with no password gate and serves a confidential document at 200.

Style and grammar governance Jul 31 – Aug 1 · 2 sessions · 520 min

Diagnosed why a technical document reached Limore unreadable and found three independent causes. The deciding one: grammar-gate.py globs **/*.html only, so all markdown was ungated by construction. Nothing had been bypassed, because no rule existed.

  • projects/shur/report-grammar/STYLE-REGISTRY.md written as the artifact-type-to-style-guide mapping, with the rule added to CLAUDE.md at both global and project level
  • Three-grammar comparison published: the same two documents in house voice, Google developer style, and ASD-STE100
  • style_guide: and reader: approved for the schema on 2026-08-01, written into SCHEMA-REFERENCE.md, resolved by a new router at system/tools/style-router/style_router.py, and applied to 1,670 vault documents
  • Verified additive frontmatter cannot break a .base: a base reads only the properties named in its own columns and filters, so a new property is invisible to every existing view. Checked across all 42.
The dry run beat the careful rule. The first routing pass sent 197 client documents to the technical style guide, because type: strategy_doc covers both a sprint plan and a client board brief. Reading the sample caught it; reading the count would not have.

Dev sprint 1 and the feature register Jul 27 – 29 · 5 sessions · 1,910 min

  • Harness right-sized for Opus 5 and Fable: a 72% cut with three adversarial verification passes before install. Round 1 found 8 drops, round 2 found 10 blockers, round 3 found 2; roughly half of round 2 was damage caused while fixing round 1.
  • Method primitive derived from 26 real builds. It could not be written from the notes because it is not in them; it had to be traced from the artifacts where it was performed. The editorial contract turned out to be the real invariant across all 26 and had never been named as part of the method.
  • shuriq-lab GitHub org stood up with its access configuration, alongside the dashboard delivery grammar (Route = Channel x Form x Reader x Action, 16 sections) and Channel and Form layers added to the feature registry at 32 rows
  • Nine-agent council decided StoryLine over Storyteller Suite, 36 to 13, with the adversary narrowing scope rather than flipping the pick
  • 5,215 lines of sprint specification audited and cut, then the sprint rebuilt from KG's and Alex's actual designs into a 103-row feature register with a Google Doc and Sheet for team review
  • Every sprint artifact reframed around what exists and what to build. Jonny rejected the first version on framing alone; the same 120 findings read as a dysfunctional team or as a build plan depending only on how they were written.
  • Claude Design pathway package delivered with a source-traceable DBM Global data extract carrying file:line citations, plus the sprint plan pass covering monorepo layout, GitHub org options, Jira process, CI gates, and work items WI-01 through WI-13
  • Command Center task views wired into TaskNotes/Views/command-center.base; July operating expense invoice produced

Harness lab: Osaurus, Kimi K3, Telegram relay Jul 27 – Aug 2 · 3 sessions · 475 min

The local-tier question settled. Five agent evaluations against the Osaurus harness on 2026-07-27: a 26B local model ignores a multi-phase research prompt entirely while producing fluent uncited prose, and Kimi K3 holds the same prompt unchanged. Osaurus raised from hold to keep. Four agent configs written, two run. ~/.dotfiles went under version control, closing the gate-zero risk from the 2026-07-25 handoff.

The Kimi K3 interactive 401 found and fixed on 2026-07-31. The key sat in customApiKeyResponses.rejected in ~/.claude.json, so interactive sessions silently fell back to Max OAuth while headless -p runs still used it. Key moved to approved, base URL trailing slash aligned, five-point acceptance test closed.

K3 proven on real work on 2026-08-02: the Black Flag v3 open-items trial ran fully isolated on experiment/kimi-k3-blackflag with both gates green. A Telegram relay was built and proven live as K3's remote channel, after Claude Code Remote Control turned out to work only against api.anthropic.com. Three first-run defects were fixed the same day.

Weekly review pipeline Jul 31 · 1 session · 30 min

  • W30 generated from 4 session records, 31 commits and 6 daily notes; the W30 editorial site built; the 13-entry _deploy archive reassembled; deployed to weekly-reviews.pages.dev on the weareshur account and curl-verified 200 on root, /W30/, /W29/, and a legacy page
  • Recorded the staging hazard: _deploy/2026-05-12-build-report and _deploy/2026-06-05-grammar-v07 exist only in _deploy and must be backed up before any rm -rf. Honoured again in this run.

Carrying Forward

Black Flag pitch

Style and grammar

Harness lab

Blocked

No overdue table this week. TaskNotes was retired 2026-07-25 and reinstated 2026-07-28, so the week carries no continuous task state to age items against. The 2026-07-30 daily note holds four hand-written tasks in plain markdown; three of the four are still open and appear above.

Project Activity

ThreadSessionsMinutesDeploysNotes
Black Flag pitch51,80553 versions, 3 megaeval rounds, 1 leak closed
Dev sprint 151,9100103-row feature register, shuriq-lab org
Style and grammar252051,670 documents restyled, 2 schema fields
Harness lab34750Osaurus to keep, K3 live, Telegram relay
Weekly reviews1301W30 shipped and verified

Each session record is counted once, against its dominant subject. The minutes column sums to 4,740 and the deploy column to 11.

Next Week Priorities

  1. Extend grammar-gate.py to markdown (WI-15) The schema fields landed 2026-08-01 and unblocked it. Until it ships, every markdown document in the vault names a style guide that nothing checks.
  2. Fix the stale decay figures at source graph/sbpi-stack-rank.md publishes numbers the live pipeline contradicts. Corrected text is already written on the experiment branch.
  3. Close the Black Flag bios and the data room Section 11 needs the remaining bios; the data room needs Jonny to pick shareable pages from the technology PDFs.
  4. Decide the 52.5 composite In the film sequences. One-number edit and a two-minute re-render if it changes.
  5. Push or merge experiment/kimi-k3-blackflag It is local only and holds both the K3 trial and the figure fix.
  6. Repair the two files with broken YAML frontmatter Confirmed pre-existing at HEAD, not caused by the 2026-08-01 bulk write.
  7. Point the comparison site's Slack panel at the standalone post And remove the duplicate copy it currently carries.

Maintenance Actions

Insights

Patterns

  • Seven days of activity, the first unbroken week in the record. The two longest sessions (18 hours and 13.3 hours) both ran overnight across a date boundary.
  • Three of the five threads produced the same finding in different form: something believed to be governed was not. grammar-gate.py checked zero markdown files, the shuriq-kg three-store portability claim compared different SPARQL semantics rather than different stores, and two pages behind a password gate were serving publicly at 200.
  • Correctness came from the verification layer, not the building layer. The harness refactor took three adversarial passes, and the megaeval's own round 3 asserted a BLOCK finding that the source table disproved.

What slowed progress

  • The deck was built three times because a pitch and a report are opposing forms: the report is built to state what is weak, the deck to state what is strong. One pipeline asked for both produced three rejected cuts.
  • Framing cost a full rebuild. Jonny rejected the sprint artifacts on framing alone, and the same 120 findings had to be rewritten as a build plan.
  • Cloudflare Pages branch aliases lag ten to twenty seconds behind a deploy, which read as a broken gate three separate times before it was recognised as propagation.
  • The deploy-guard hook infers its scan target from the shell's working directory, so chaining a cd from the vault root made it grammar-gate the entire vault and block the deploy.

What went well

  • The dry run beat the careful rule: 197 misrouted client documents were caught by reading the sample, which reading the count would have missed.
  • Failures were proved pre-existing before being claimed. Two files failed a YAML re-parse after the bulk write, and parsing the same two at HEAD showed both were already broken.
  • Read-only auditors returning structured findings, with the main loop applying a triaged ledger, gave the megaeval full coverage without ever letting an agent write into the build tree.
  • The megaeval gained a rejection layer. CORRECTIONS-mandated statements, film vocabulary and the intake's own words outrank register preferences, and about a third of auditor fixes were rejected on those grounds.

Needs attention

  • A PASS from a gate that examined zero matching files is indistinguishable from a real PASS. One confirmed instance, and no detection for the next one.
  • Working from the artifacts people made beats working from the specification about them. The route grammar was written by Claude sessions and presented as settled; Jonny could not use it, and the document it came from was only ever a demo device.
  • Agents write numbers they have not counted, including numbers supplied to them in a prompt. Every invented count in the 2026-07-27 session came from an agent stating a figure it had not measured.
  • A 157-word explanation of the report engine did not reach Limore, which meant it was never going to reach KG either. The version that worked was 72 words and led with what he had to send.