Skip to content

[AIGTWY-4880] CUJ5: Verify budget recommendations and agent startup - #911

Merged
andy-xu-db merged 25 commits into
mainfrom
andy-xu-db/aigtwy-4880-budget-defaults
Oct 6, 2026
Merged

andy-xu-db merged 25 commits into
mainfrom
andy-xu-db/aigtwy-4880-budget-defaults

Conversation

@andy-xu-db

@andy-xu-db andy-xu-db commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add CUJ5 coverage for budget recommendations and agent startup in AIGTWY-4880. Two service principals use the same published configuration with one fixed 1% tier (spending_percentage: 0.01): the low-spend principal keeps the Claude/Sonnet default, while the above-tier principal receives a Codex/Luna recommendation.

Coverage

  • Low-spend principal: verify the real managed config and Claude/Sonnet recommendation, launch bare ug, and check the generated model setting and native Claude Sonnet 4.6 header.
  • Above-tier principal: verify the published Sol default, compare ug usage spend, threshold, and percentage against backend reads before and after the command, allowing concurrent spend increases. Verify the real Codex/Luna recommendation, launch bare ug, check that the launch output displays that recommendation, and verify Codex reaches its interactive prompt and exits normally.

The above-tier test covers the recommendation and Codex startup. It does not assert that Codex applies Luna or overrides its managed Sol default.

A fresh principal may have no spend counter. When budget figures are available, the low-spend case requires usage below 1%; when both figures are absent, it verifies default model selection without claiming a measured spend percentage. Neither test submits an inference task or writes the budget. Exact boundary checks, multiple-tier precedence, and explicit-agent overrides are outside this coverage.

Test setup and CI

Both CUJ5 cases run alongside every other dedicated-workspace CUJ in the shared E2E CUJs job, through one pytest collection. There is no separate CUJ5 job, matrix, marker filter, or cross-run concurrency lock.

Case GitHub secrets
Above-tier / existing CUJs UG_CUJ_SP_CLIENT_ID, UG_CUJ_SP_CLIENT_SECRET
Below-tier default UG_BUDGET_CUJ_SP_CLIENT_ID, UG_BUDGET_CUJ_SP_CLIENT_SECRET

The workflow exposes both credential pairs. Each test class selects its own pair through the shared base fixture and gets a separate SDK client, temporary home, and artifacts. Local agent configuration is cleaned up between classes through ug revert; workspace and budget configuration remain read-only.

The runner installs both agents because the managed configuration enables both. Pytest runs directly with --confcutdir=tests/e2e_cuj. The single job contributes to the existing integration gate. No account-level authentication is required.

Validation

  • 50 focused CUJ scaffold and existing CI-contract tests passed after cleanup. All 14 CUJ cases collect together without a dedicated pytest configuration or marker filter; Ruff lint/format, type checks, and diff checks passed.
  • The combined E2E CUJs job passed: 12 passed, 2 existing smart-routing skips, no deselections. Both CUJ5 identities passed after the smart-routing CUJ, validating cleanup between classes.
  • No production changes to src/ucode/agents/codex.py are included.

@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from 395c82d to fbbe18b Compare September 30, 2026 19:55
@andy-xu-db andy-xu-db changed the title [AIGTWY-4880] Add budget-driven smart defaults E2E tests [AIGTWY-4880] Exercise budget smart defaults in live E2E journeys Sep 30, 2026
@andy-xu-db
andy-xu-db changed the base branch from main to andy-xu-db/stack/andy/aigtwy-4880-budget-evidence September 30, 2026 19:56
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from fbbe18b to 2895e48 Compare September 30, 2026 20:16
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch 2 times, most recently from 3757625 to a9beae9 Compare September 30, 2026 20:26
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from 2895e48 to b9aa6ee Compare September 30, 2026 20:26
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from a9beae9 to edf1288 Compare September 30, 2026 20:39
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch 2 times, most recently from 2a9ea3a to fac2cbe Compare September 30, 2026 20:49
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from edf1288 to caddcd4 Compare September 30, 2026 20:49
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from fac2cbe to f6e7b92 Compare October 1, 2026 14:38
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from caddcd4 to 4ff6751 Compare October 1, 2026 14:38
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from f6e7b92 to 8deffc9 Compare October 1, 2026 15:05
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch 2 times, most recently from 989bf6d to bf245ba Compare October 1, 2026 15:09
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch 2 times, most recently from d24aed3 to fbe142c Compare October 1, 2026 16:24
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from bf245ba to bf3e3ac Compare October 1, 2026 16:24
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from fbe142c to d45bc02 Compare October 1, 2026 16:29
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from bf3e3ac to d1505fc Compare October 1, 2026 16:29
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from d45bc02 to f835a8e Compare October 1, 2026 16:37
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch from d1505fc to 5e656bd Compare October 1, 2026 16:37
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from f835a8e to e5fe6dc Compare October 1, 2026 17:17
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/stack/andy/aigtwy-4880-budget-evidence branch 2 times, most recently from 715d107 to b890d33 Compare October 1, 2026 18:51
@andy-xu-db
andy-xu-db force-pushed the andy-xu-db/aigtwy-4880-budget-defaults branch from e5fe6dc to bc1dab6 Compare October 1, 2026 18:51
Comment thread .github/workflows/ci.yml Outdated
@andy-xu-db

Copy link
Copy Markdown
Collaborator Author

This PR contains no writes. We do not need to update the budget threshold at all. The service principal for below tier tests never makes any inference requests so its spend will not increase. The service principal for above tier tests will always be above the current spend tier.

@david-siqi-liu david-siqi-liu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. One P1 to follow up on, not blocking:

P1

  • The above-tier journey proves the recommendation through the startup banner (Recommended agent is Codex with model ...) and exits without completing a turn, and the below-tier case relies on settings plus a startup label. AGENTS.md: "neither routing banners nor native records alone prove an applied decision." Completing one task through the recorder and checking that the inference request model and the native turn model (SessionEvidence, canonical_model) are Luna, and Sonnet below the tier, would show the recommendation is actually applied.

FYI, the terminal ug revert you added to the cuj fixture teardown also covers the cleanup problem breaking #965 and #990 in CI. It's class scoped, so those PRs still need their own handling within a class.

@david-siqi-liu

Copy link
Copy Markdown
Collaborator

Follow-up after checking the test against the CUJ 5 plan in the UG Configure E2E doc. My approval stands, but I think these are must-fix before this counts as CUJ 5 coverage:

P0 (coverage)

  • The fixture has a single 1% tier (Codex/Luna), so tier selection, which is the core of CUJ 5, is mostly untested. Not exercised: Codex on Sol between 50% and 80%, Luna above 80% with the highest applicable tier replacing Codex's Sol default, the exact 50%/80% boundaries (the plan allows marking these unverified if spend can't be held stable), ug claude while a tier recommends Codex (Claude on its own default with the recommendation shown), and a Codex-only recommended model never being applied to Claude.
  • The plan asks each launch to verify the actual agent process, request model, and task completion. Neither test completes a task or checks an inference request model (the banner-only point from my earlier comment), so the launch side of the plan isn't covered yet.

The PR body marks the boundaries and explicit-agent overrides out of scope; flagging them here so they're tracked against the plan rather than dropped.

@andy-xu-db

Copy link
Copy Markdown
Collaborator Author

UG Configure E2E doc

I discussed this with Lilly previously. Since these tests require a separate service principal for each possible case, we decided to test with only one spend tier which is enough to prove that the spend tier recommendations work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ug-review Run the automated UG review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants