Repository navigation
[AIGTWY-4880] CUJ5: Verify budget recommendations and agent startup - #911
Conversation
395c82d to
fbbe18b
Compare
fbbe18b to
2895e48
Compare
3757625 to
a9beae9
Compare
2895e48 to
b9aa6ee
Compare
a9beae9 to
edf1288
Compare
2a9ea3a to
fac2cbe
Compare
edf1288 to
caddcd4
Compare
fac2cbe to
f6e7b92
Compare
caddcd4 to
4ff6751
Compare
f6e7b92 to
8deffc9
Compare
989bf6d to
bf245ba
Compare
d24aed3 to
fbe142c
Compare
bf245ba to
bf3e3ac
Compare
fbe142c to
d45bc02
Compare
bf3e3ac to
d1505fc
Compare
d45bc02 to
f835a8e
Compare
d1505fc to
5e656bd
Compare
f835a8e to
e5fe6dc
Compare
715d107 to
b890d33
Compare
e5fe6dc to
bc1dab6
Compare
This reverts commit 1b04347.
…faults' into stack/andy/aigtwy-4880-budget-defaults
|
This PR contains no writes. We do not need to update the budget threshold at all. The service principal for below tier tests never makes any inference requests so its spend will not increase. The service principal for above tier tests will always be above the current spend tier. |
david-siqi-liu
left a comment
There was a problem hiding this comment.
Approving. One P1 to follow up on, not blocking:
P1
- The above-tier journey proves the recommendation through the startup banner (
Recommended agent is Codex with model ...) and exits without completing a turn, and the below-tier case relies on settings plus a startup label. AGENTS.md: "neither routing banners nor native records alone prove an applied decision." Completing one task through the recorder and checking that the inference requestmodeland the native turn model (SessionEvidence,canonical_model) are Luna, and Sonnet below the tier, would show the recommendation is actually applied.
FYI, the terminal ug revert you added to the cuj fixture teardown also covers the cleanup problem breaking #965 and #990 in CI. It's class scoped, so those PRs still need their own handling within a class.
|
Follow-up after checking the test against the CUJ 5 plan in the UG Configure E2E doc. My approval stands, but I think these are must-fix before this counts as CUJ 5 coverage: P0 (coverage)
The PR body marks the boundaries and explicit-agent overrides out of scope; flagging them here so they're tracked against the plan rather than dropped. |
I discussed this with Lilly previously. Since these tests require a separate service principal for each possible case, we decided to test with only one spend tier which is enough to prove that the spend tier recommendations work. |
Summary
Add CUJ5 coverage for budget recommendations and agent startup in AIGTWY-4880. Two service principals use the same published configuration with one fixed 1% tier (
spending_percentage: 0.01): the low-spend principal keeps the Claude/Sonnet default, while the above-tier principal receives a Codex/Luna recommendation.Coverage
ug, and check the generated model setting and native Claude Sonnet 4.6 header.ug usagespend, threshold, and percentage against backend reads before and after the command, allowing concurrent spend increases. Verify the real Codex/Luna recommendation, launch bareug, check that the launch output displays that recommendation, and verify Codex reaches its interactive prompt and exits normally.The above-tier test covers the recommendation and Codex startup. It does not assert that Codex applies Luna or overrides its managed Sol default.
A fresh principal may have no spend counter. When budget figures are available, the low-spend case requires usage below 1%; when both figures are absent, it verifies default model selection without claiming a measured spend percentage. Neither test submits an inference task or writes the budget. Exact boundary checks, multiple-tier precedence, and explicit-agent overrides are outside this coverage.
Test setup and CI
Both CUJ5 cases run alongside every other dedicated-workspace CUJ in the shared
E2E CUJsjob, through one pytest collection. There is no separate CUJ5 job, matrix, marker filter, or cross-run concurrency lock.UG_CUJ_SP_CLIENT_ID,UG_CUJ_SP_CLIENT_SECRETUG_BUDGET_CUJ_SP_CLIENT_ID,UG_BUDGET_CUJ_SP_CLIENT_SECRETThe workflow exposes both credential pairs. Each test class selects its own pair through the shared base fixture and gets a separate SDK client, temporary home, and artifacts. Local agent configuration is cleaned up between classes through
ug revert; workspace and budget configuration remain read-only.The runner installs both agents because the managed configuration enables both. Pytest runs directly with
--confcutdir=tests/e2e_cuj. The single job contributes to the existing integration gate. No account-level authentication is required.Validation
src/ucode/agents/codex.pyare included.