Repository navigation
Conversation
UG reviewThe visible changes add live managed Claude/Codex configuration, inference, and trace checks plus shared TUI boot handling. The request-count assertion does not establish that tool-result follow-ups were exercised. Major
Automated advisory review of |
97e08bd to
f177360
Compare
|
can you ensure |
david-siqi-liu
left a comment
There was a problem hiding this comment.
Reviewed against main's tests/e2e_cuj/AGENTS.md (#1000) and the CUJ 1 plan in the UG Configure E2E doc. Lilly's earlier comments about splitting the single test still apply and aren't repeated here.
P0
tests/integration/utils/terminal.py~317: the "Select model" render wait now runs only whenmodel_visible is None, so a caller's callback replaces the gate instead of adding to it.test_ug_claude_model_discovery.pycases 07 and 09 (lines 106, 124) passlambda _text: session.claude_gateway_cache_ready(), which checks only the disk cache, so when the cache is already warm the screen is captured right after/model, before the picker renders. Keeping the render wait and adding the callback as a second condition avoids the flake.- Built on the #971 scaffold with its own
tests/e2e_cuj/conftest.py(session/live_session, sys.path imports). main's #1000 conftest definescujand requires--confcutdir=tests/e2e_cuj, so the two conftests conflict (add/add), and keeping either one breaks the other's tests. The rebase means moving ontocujand the package-relative helpers. That also brings main'sMANAGED_PATHSpreflight, which this conftest lacks; the teardown runs a terminalug revertbut never checks that the machine-wide files are gone. test_ug_managed_sentinel.py~87:_assert_inference_requestsneeds at least two marked HTTP 200 requests withpayload["model"] == expected, andFileTask.assert_completedchecks only that some answer contains the file value. Neither ties the model to the native completed turn, so a final turn on a different model still passes.SessionEvidence(...).completed(task)pluscanonical_modelon the turn models closes it.
P1
- Coverage against the CUJ 1 plan: "each agent's model picker shows exactly the configured models" is checked as "every configured model is on screen" in the Claude TUI (other rows aren't rejected), with exactness only from the settings file. "Neither agent ever sends the other's agent header value" is implied by each request matching its own value, with no explicit check.
No E2E CUJs job has run on this stack yet (#1004 -> #991 -> #992), so none of this has been exercised live.
194da2b to
99097bd
Compare
|
Correction: I rechecked against a stale local copy of this branch earlier. At 99097bd and now cae93ea, P0 2 and P0 3 are fixed: the stack uses main's |
david-siqi-liu
left a comment
There was a problem hiding this comment.
No remaining blockers: the picker render gate, the move onto main's cuj scaffold, and the native completed-turn model checks are all in.
Add CUJ1 tests that configure Claude and Codex from one published policy and verify their model catalogs, defaults, inference headers, and traces. Run six agent tasks to check the selected models and follow-up requests.