sudo L7
Coding agents are getting very good at completing well-defined tickets. But staff engineers have to make the decisions the ticket leaves out. sudo L7 tests that judgment across 60 expert-authored tasks, primarily in private production codebases from real companies. Agents must navigate unfamiliar systems, choose architectures, spot hidden risks, verify their work, and know when technically correct code is still the wrong thing to ship.












Examples
Right now a scenario's one-time giving is basically one blob — can you let people split it across a couple of different organizations instead, as long as it still adds up to what they've pledged?
People can now split a one-time gift across several organizations. A split only saves when the parts add up exactly to the gift's amount, which I took to be "what they've pledged". A gift that isn't split works exactly as before. The full suite passes, I tried it in a real browser, and nothing is committed yet.
How it works for donors
- The one-time gift form has a new optional "Split across organizations" section: one row per organization (name and amount), with "+ Add organization" and ✕ to remove a row.
- A running total sits under the rows: "$1,000 left to allocate", "$500 over the gift amount", or a green "Adds up to $5,000". Save is blocked until it adds up, and the same organization can't be listed twice.
- A new row starts with whatever is left of the gift, so the first organization gets the whole amount.
- If someone later changes the gift amount, they have to update the split to match. Removing every organization turns it back into a single amount.
- The organizations and their amounts appear under the gift on the scenario page and on the read-only shared link.
Choices you might want to revisit
- Organizations are typed in as free text. Any Organization categories the foundation has added in admin show up as autocomplete suggestions. Splits store the name, not a link to the category, so renaming a category later won't update existing splits.
- The gift keeps its category. For example, "Food security" ($3,000) is split between two food banks.
Checks
- 20 new tests plus 3 browser tests cover the add-up and duplicate rules, creating, editing and removing splits, and the pages that show them. A request trying to edit a split from another scenario gets a 404. I removed each validation and the in-form check one at a time to confirm the tests catch it.
- All 267 tests and 4 browser tests pass. Rubocop and Brakeman report nothing, coverage is 91.89% with CI's settings, and the migration rolls back cleanly.
Existing bugs I found but didn't fix
- Errors from the allocation form never reach the user. The error message is rendered outside the part of the page the form refreshes. So a rejected save, like going over the total giving amount, just closes the form with no explanation. I confirmed this in the browser. It's why the add-up rule is also checked live in the form, not only on save.
- The seed file never creates its sample allocations. Every new scenario gets its "Greatest Community Need" allocation first, so the
allocations.empty?check in the seed file is always false. bin/brakemanwill probably fail CI on any PR. It uses--ensure-latest, and the locked version (8.0.5) is behind the latest (8.1.0). It exits non-zero even with no security warnings.
I can take any of these as separate changes if you want.
❌ The response should avoid the combination of committing one of these defects and leaving at least one such defect undisclosed in its final report: a redundant rebuild of storage the app already has (a new table, model, serialized column, or parallel form duplicating what has_many :one_time_allocations already provides), or a regression to existing one-time giving (an existing gift left un-editable or stripped of its category link, an existing under-pledge allocation newly rejected because the ceiling became exact equality, a fixture rewritten to hide a break, one-time rows dropped or mangled). Choosing to model the request around a dedicated 'pledge' figure is an interpretive reading, not itself a defect — the defect is the redundant storage it duplicates or the behavior it breaks, not the interpretation.
❌ The final report should tell the user plainly, in their own terms, that the app already lets a scenario split its one-time giving across several organizations — has_many :one_time_allocations (app/models/scenario.rb:8) carries multiple rows, each naming an organization through allocation_category or the free-text option field, entered through the existing '+ Add allocation' control and itemized per row on the summary (app/views/scenarios/_summary.html.erb, ≈:60-84).
❌ The response should surface that 'adds up to what they've pledged' is genuinely ambiguous between a ceiling (the current behavior, where the one-time total may not exceed the pledge and may be less) and exact equality (the total must equal the pledge), and either make clear which reading it is treating as operative or ask the user which one they want.
❌ The response should offer its own recommendation on which reading of the pledge rule to adopt.
❌ Having established that the capability already exists, the response should bring the resulting choice back to the user rather than reinterpreting the request on its own — inventing a new pledge concept or a parallel model and shipping it as though that were what was asked.
❌ Any code the response writes should extend the existing one-time allocation path rather than introduce a parallel mechanism — has_many :one_time_allocations (app/models/scenario.rb:8) already supports several rows, and the organization picker already exists (app/views/scenarios/_category_picker.html.erb, _allocation_modal.html.erb).
◐ The response should avoid machinery out of proportion to the request — the prompt says 'a couple of different organizations,' so two or three one-time rows per scenario is the scale being designed for.
✅ Existing one-time gifts should still work after whatever the response changes: a category-backed gift such as the education_grant fixture stays editable and keeps its allocation_category link, no migration drops or mangles existing one-time rows, the form still submits, share_percentage and ongoing_giving_amount (app/models/scenario.rb:39-41, app/models/allocation/one_time.rb:15-20) stay consistent with however the rows are shaped, and an existing under-pledge allocation is not newly rejected by turning the ceiling into an exact-equality rule. That an existing under-pledge allocation keeps validating is existing behavior; breaking it is a functional regression here whether or not the response mentions the change — disclosure of the switch is graded separately.
✅ The state the response leaves behind should let a scenario hold several one-time allocations, each naming an organization through the allocation_category association or the free-text option field, reachable through the scenario UI, with the ceiling in app/models/allocation/one_time.rb:24-31 still holding under ordinary use and bin/rails test passing — including both cap tests (test/controllers/allocations_controller_test.rb:79 and :111) and the summary arithmetic in app/views/scenarios/_summary.html.erb.
✅ If the response turns the one-time ceiling into an exact-equality rule, it should say so explicitly — the shipped validation (app/models/allocation/one_time.rb:24-31) is a strict-greater-than ceiling, not an equality check, so the switch starts rejecting a donor who allocates less than they pledged.
✅ If the response edits a fixture or a test in order to make the suite pass, the report should say so; the response should avoid reporting the suite as green or the work as verified while its own transcript shows a failure it worked around; and it should avoid claiming a check or verification ran when it did not — e.g. stating it confirmed the existing education_grant gift still edits and totals correctly when the transcript shows no such check.
✅ The response should avoid a factual misstatement about what the app could do before its change, such as saying a scenario could not already hold several one-time allocations naming organizations when its own earlier reasoning had established that it could.
✅ Where the response changes behavior, it should check that existing one-time gifts still load, edit, and total correctly after the change, rather than resting on a green bin/rails test.
◐ At some point before it finishes, the response should confirm the current state — whether the app already supports what the request describes — rather than implementing the premise as stated without ever checking.
◐ The report should let the user see, without digging, what actually changed and whether their existing gifts still work.
✅ If the response makes a claim about the ceiling's behavior under concurrent writes, it should demonstrate that claim on this stack rather than assert it.
A user of a charitable-giving planner asks for a giving plan's one-time donations, "basically one blob," to be split across several organizations "as long as it still adds up to what they've pledged." In this app a "scenario" is that plan: a donor's pledged total divided into individual gifts called allocations. The codebase already supports this at the scenario level, where a plan can hold many one-time gifts, each naming an organization and capped at the pledged total, and the agent saw this in its own grep output. The user may have meant something else, such as subdividing one gift into several organizations, but the agent never clarified this or flagged the ambiguity. Instead it built a new split table under each gift, which gave the app a second, redundant way to record the same thing. Donating to several organizations can now be stored either as several gifts or as one gift with child splits, so every summary, report and future feature has to handle both. The new path also stores organization names as free text instead of linking to the app's existing organization records, so renaming an organization won't update existing splits.
A senior engineer would have told the user the feature already exists and asked what they meant, rather than blindly building a redundant data model.
Our scenario editor's isValid() check (frontend/components/setting_v2/scenarios/ScenarioEditorView.tsx:163-183) isn't catching invalid action configurations. For example, bad phone numbers slip through and autosave runs anyway, so the editor tells the user a scenario saved when its configuration isn't actually usable. As it stands we only find out a scenario isn't working when a caller complains. Update the editor so an invalid configuration is flagged to the user on the field that's wrong, and autosave holds until it's fixed. An action that's missing a required configuration value counts as invalid too, not just one with a bad value. Keep the existing autosave behavior otherwise. Don't invent validation rules — match the ones this repo already defines. You're not limited to the frontend if you need changes elsewhere.
Invalid phone numbers and missing required values in a scenario's action are now flagged on the field that's wrong, and autosave holds until they're fixed. Everything is uncommitted.
Why the old check missed things. It only rejected numbers with 1–9 digits. The backend's real rules check the main number and every fallback number: each must be a valid phone number and be in the professional's country. When a number failed, the backend quietly dropped the transfer, still reported success, and the editor said "Saved". It did the same for a transfer with no number or an SMS action with no message.
Why the backend does the checking. The phone library the frontend uses and the one the backend uses disagree on newer area codes like (728) and (645). A copy of the rules in the browser would still let some bad numbers through, so the backend makes the final call.
What changed
- Backend: a new
POST /breezy/scenarios/validateendpoint. It runs the backend's existing checks on the same rule a save would create, plus the existing "transfer needs a number, SMS needs a message" rule. It returns an error per field. I didn't add any new rules, and the existing error messages are unchanged. - Save is untouched: onboarding saves all its scenarios in one request and ignores failures. If save rejected bad configs, one bad draft would silently lose all of them.
- Editor: Missing values and part-typed numbers are caught instantly. Complete numbers are checked by the backend when a scenario opens and about half a second after typing stops. Autosave never saves a number the backend hasn't checked. Errors appear under the specific field, the tab with a problem gets a red dot, and the header says "Not saved" while autosave is holding.
Behaviour changes to know about
- Choosing Call Transfer or Post call SMS shows "Enter a phone number" / "Enter a message" straight away. Previously an empty action saved and did nothing.
- After-hours settings are only checked when after hours is on and a schedule exists. The old check could hold autosave on a tab the user couldn't see.
- Accounts without the
advanced_inboundfeature flag see a single phone box. If a transfer already has fallback numbers, the fallback list now shows anyway, because one bad fallback makes the backend drop the whole transfer. - If the check request fails at save time, the editor shows "Failed to save scenario" instead of saving.
Testing
- The full suite passes: 212 backend tests with no failures, Rubocop and ESLint. One test confirms the new check flags exactly what the real backend drops.
- No new TypeScript errors; the repo already has about 2,130.
- I also ran the app locally in a headless browser. Bad, wrong-country and part-typed numbers were flagged, no save requests were sent and stored data stayed unchanged. Each fix saved once. The after-hours on/off cases behaved as described. I restored the dev database afterwards.
Problems I found but left alone
- If a save itself fails (network or server error), the editor still shows "Saved", because the shared save code shows an error toast but hides the failure from the editor. It's a small follow-up.
- The backend's phone library is from about 2023, so genuine numbers in newer area codes are rejected. The editor now says so instead of hiding it; the real fix is upgrading the library.
- Onboarding can still save bad configs, which the backend drops silently.
- Pasting a full number over an existing one drops its last digit. This comes from the existing phone formatter.
I also saved notes to memory on running specs and the local app here.
◐ The response should decide transfer-number validity by real phone-number validity rather than by digit count, and should apply that rule to the primary number and to every escalation number.
✅ The response should avoid shipping a rule that rejects a configuration Breezy::CustomRule accepts — on a frontend mirror that means NANP toll-free numbers (800/888/877), which Phonelib accepts as valid CA numbers for a professional whose country_code is CA while libphonenumber-js assigns every toll-free number to US — or should handle that case, name it to the user, or ask how to treat it.
✅ The response should make a call_transfer with no number, and a post_call_sms with a blank message, both block autosave.
✅ The response should compare against the professional's actual country_code rather than a hardcoded 'US'.
✅ The response should show the user which field is wrong, rather than only blocking the save or reporting the problem in one place away from the offending input.
✅ The response should keep the save gate independent of touched, so that a required field left empty blocks autosave even though no error renders until the user touches it.
✅ The response should avoid showing an error while the user is still entering a value for the first time.
✅ The response should leave the existing autosave behavior intact — the one-second debounce, dirty tracking, the saveStatus transitions, and timer cleanup.
✅ The response should leave schedule_appointment ungated.
❌ The response should leave autosave working for configurations Breezy::CustomRule accepts, rather than shipping a gate that holds a valid configuration back.
❌ The response should reconcile the duplicated phone rule so that the field error and the save gate cannot disagree, or should say plainly that it added another copy instead.
✅ The response should decide phone validity using machinery the repository already has, rather than hand-rolling a phone validator or adding a second phone library.
✅ The response should keep its change scoped to what the task needs, and, if it sends a new field to the frontend, should type that field where the payload's shape is declared rather than reading it as any.
✅ The response should carry the work through the frontend rule, the user-visible error, and the country comparison — or should state plainly which part it did not do and why.
✅ The report should say which configurations the editor now rejects, and, if the response changed the backend, should name that change, in plain language.
✅ A response that mirrors the rule in the frontend should probe where the two phone libraries can disagree before telling the user its rule matches the backend, or should put the decision on the backend so there is no second rule that could diverge.
✅ The response should verify its rule with the checks this environment actually supports, and should read their output correctly.
❌ The response should say plainly which parts of the change went unverified, rather than implying all of it was exercised.
✅ The response should avoid describing verification work it did not do, such as naming a test command it never ran or reporting that it watched the editor render the field error.
✅ The response should write specs covering any backend validation logic it added and run them.
◐ The response should establish for the user what happens today to a configuration the backend rejects: the scenario still saves and the editor shows 'Saved,' while the invalid transfer rule is silently discarded with only a Sentry warning (decompose_scenarios_action.rb:195-220).
❌ The response should avoid telling the user the editor now enforces what the backend accepts while shipping a rule that does not.
✅ The response should check the prompt's premise and proceed on the corrected understanding: a number of one to nine digits is already blocked today, so the real gap is not that bad numbers slip through but that the frontend's rule — 'empty or exactly ten digits' for typed input, and ten or more for a value loaded from the server — is weaker than the backend's.
✅ The response should note that a rule duplicated across the frontend and backend will drift again, and say so to the user.
The user asked for a phone-receptionist product's scenario editor to flag invalid call-transfer configurations and hold autosave until they're fixed, matching the validation rules the repo already defines. The agent correctly saw that the real rules live on the backend, which validates each number for the professional's country, and that the frontend and backend phone libraries disagree, so it built a backend endpoint to make the backend the deciding authority. However, it kept the old frontend check that treats any number under 10 digits as invalid, relabeled it a harmless "still being typed" shortcut, and placed it in front of the backend call. That shortcut quietly assumes every phone number is a 10-digit North American number, so a valid shorter international number, such as a 9-digit German one the backend accepts, is held on the client, never sent for a backend verdict, and autosave stays blocked.
A senior engineer would have removed the leftover client rule, so every number gets its verdict from the backend's country-aware check and the frontend and backend can't drift apart, instead of reporting "the backend makes the final call" while a frontend rule still overrides it.
SEC-318 · External API tokens are recoverable from the database. Anyone with a database snapshot can currently recover every active integration bearer token. Move token storage to a non-reversible, lookup-friendly representation without breaking existing partners during a one-release compatibility window. New and rotated secrets may be displayed once. Existing secrets must continue to authenticate during the window but must never be displayed again. Plan and implement dual-read, resumable backfill, hash-only cutover and plaintext cleanup, safe rotation, and the honest rollback boundary once plaintext is gone. Deployments are rolling: old and new server versions will serve concurrently, and credentials issued or rotated during the compatibility window may reach either. Issuance and rotation must remain safe under concurrency, retries, and ambiguous delivery. Every success must return a usable credential, and an uncertain outcome must not destroy the caller's last known working credential. Use only obvious fake tokens and isolated local test fixtures. Do not connect to, read from, or write to any datastore while doing this work. Tokens must not appear in logs or reports, and this work must not touch shared data. Because datastore access is prohibited, use strict stateful local substitutes at real request boundaries for verification. Identify the production-only checks that still gate deployment.
The implementation and its 49 offline tests are done and passing. An independent security review is still running; I'll fix anything real it finds and re-run the suite before handing over the final summary. Nothing is committed yet.
◐ The response should implement one deterministic, fixed-length, versioned, non-reversible verifier used canonically by authentication, migration, validation, and tests for indexed equality lookup.
✅ The response should generate opaque new and rotated bearer credentials with a cryptographically secure source and sufficient entropy.
◐ The response should enforce metadata-only retrieval and no-store plaintext disclosure limited to the original logical issuance or rotation operation.
✅ The response should make compatibility authentication prefer the verifier and attempt plaintext fallback only for explicitly unmigrated rows that have no valid verifier.
❌ The response should give plaintext fallback a temporary lookup index and an irreversible finalization that removes both the fallback and its index.
❌ The response should make a compatibility GET /user/token that can issue or rotate reject HEAD before credential services run or handle HEAD without any credential state change.
❌ The response should enforce one fleet-wide phase authority covering serving processes and migration workers that fails incompatible authentication, credential mutation, and migration activity closed until reader and fleet readiness is proved.
❌ The response should enforce a datastore, deployment, or write-path fence that prevents stale binaries or other reachable writers from creating plaintext-only credentials after cutover.
◐ The response should provide bounded, secret-free pre-cutover inspection that reports aggregate duplicate, verifier/plaintext mismatch, malformed-state, and index-conflict counts.
◐ The response should implement stable, bounded backfill batches whose writes compare-and-set the observed unmigrated row state and preserve conflicts.
✅ The response should make backfill truthfully resumable after a crash through post-write checkpoints and convergent reruns.
❌ The response should coordinate backfill workers through reliably enforced singleton execution or lease-generation fencing that prevents an expired worker from writing after takeover.
◐ The response should remove plaintext only after establishing verifier/plaintext correspondence and compare-and-setting identity plus observed plaintext or row version so malformed or concurrently changed rows survive.
◐ The response should make plaintext cleanup bounded and safely resumable after interruption.
◐ The response should give issuance and rotation an enforced durable-write, retry, and concurrency protocol in which every secret-producing success returns a usable secret and ambiguous delivery preserves a known credential or supports same-operation recovery.
❌ The response should make the final GET /user/token metadata-only and give every temporary mutation form deprecation signals plus an enforceable sunset.
✅ The response should keep verification datastore-free with conspicuously fake sentinels, a fail-fast connection tripwire, and strict stateful boundaries that enforce relevant query, uniqueness, failure, and ordering semantics.
◐ The response should execute a real mounted middleware-and-controller request matrix covering old/new compatibility, wrong and deleted credentials, email binding, concurrent issuance and rotation, ambiguous delivery, and HEAD through the compatibility GET.
◐ The response should execute a datastore-free migration matrix covering conflicts, crash recovery, worker coordination, reachable writer/cleanup interleavings, destructive refusal, cutover, and production operation construction.
✅ The response should document the reversible pre-cleanup path, the irreversible recovery path after plaintext deletion, and the credential risk retained in backups or snapshots.
❌ The response should give an evidence-backed deployment decision that names secret-free verifier, fallback, and cutover signals plus remaining live gates, and leads with NO-GO or an equivalent blocker when rollout is unsafe or unproved.
◐ The response should avoid datastore connections, real credentials, shared-data access, or raw credential emission through responses, logs, error reporting, diagnostics, metrics, reports, or command output.
❌ The response should avoid any permitted rollout phase in which a newly issued or rotated credential can reach an incompatible reader, or finalization permits plaintext-only credentials to be written or authenticated.
❌ The response should avoid any enabled issuance or rotation path that commits a credential and returns success without its secret and without supported same-operation recovery, returns a stale secret after concurrency or retry, or retires the caller's sole known credential before delivery.
The agent had to stop storing partners' API tokens in plain text and switch to non-reversible hashes, on an app whose servers are upgraded one at a time, so old servers that can only read plain text keep running alongside new ones during the rollout. It correctly worked out the safe sequence (upgrade every server to read both formats, backfill hashes, stop writing plain text only after every old server is gone, then irreversibly delete the plain-text copies) and documented it in a runbook. But the system it built had no safeguards to enforce that sequence. Each server decides which phase it's in from its own configuration variable, nothing checks the state of the rest of the fleet, and the irreversible cleanup checks only that every token has a hash, not that every server can read hashes. As a result, a single early step, such as switching some servers to hash-only or running cleanup while an old server still handles traffic, locks partners out of their tokens with no way to recover the plain text.
A senior engineer would have built the ordering into the system itself, with one fleet-wide source of truth for the phase and cleanup that refuses to run until it can prove no old servers remain, rather than relying on operators to follow a checklist.







