Quality & health assessments — design systems, tokens, governance, benchmarks. Scored by direct inspection, each with a prioritised action list.
search can't find agent-created documents, raw base64 reads blew the request-size cap, one empty response shipped as completed). Carries a living issue register (13 open) and an append protocol for future observation windows. Zero feedback scores exist, so every finding came from reading transcripts.agent/lawrence-2 turns out to exist twice — the chat Langfuse project's copy serves Python, the agents project's serves TS — and the copies ran different models (sonnet-5 vs sonnet-4-6) until this audit fixed it. Slot-by-slot comparison of what each runtime feeds its copy, both live templates annotated, remaining gaps ticketed (citation standards, skills catalog, loop budget, output tokens, prompt cache, content drift). Custom instructions work only on TS. Feeds LEX-874.account_env becomes, which backends the AI stack talks to, which branch it tracks, and ingestion's deploy stack). Code-verified across five codebases: three env registries, ~90 env-keyed Terraform maps, ~40 SOPS files, ~143 SSM params, 22+ vendor consoles, sorted into hard / silent / manual failure classes plus a third-party pillar. prd-gb at 1 of 32 secret files is the existence proof. Decided 23 Jul: no new env — an isolated demo firm in production instead (synthetic data out of reporting, flags + firm scoping), which the code puts at roughly three small patches plus one real payment decision (Stripe).@lawhive/ui design system. Tokens, components and adoption are strong; governance and usage docs are the two cheap, high-leverage gaps. Maturity: Managed, with the token layer already Systematic. Includes a prioritised action list (immediate / near-term / longer-term) and scope caveats.