[bot] Merge master/ebe45f35 into rel/dev - #1837
Conversation
gen-ai reports a turn that ended without a final answer as a 502 whose error event carries a `reason` (max_iterations, max_tokens, content_filter). The 502 alone routed it to the transient retry, which re-sent the same question up to five times: the work was repeated, the question landed in a conversation's history twice, and a stall was hidden from every evaluator. These reasons now raise TurnIncompleteError, a ChatError that is not retried and keeps the partial result. A 502 without a reason, or with `unknown`, is still treated as transient. This changes behaviour for every eval kind, not only conversations: an item whose turn hits the iteration limit now fails instead of being retried. jira: QA-29448 risk: low
create_metric is an upsert: on an existing id it replaces that metric and reports `created_new: false`. Cleanup listed every create_metric result as created, so a run in which the agent updated a metric that was already in the workspace deleted it. _extract_created_metric_ids now skips results with created_new false. jira: QA-29448 risk: low
A conversation passed when every turn had the expected skill active and produced an output of the expected type. Nothing compared what the output contained (the check that tried returned None for every chart in both conversation datasets), the simulated user was primed with the expected chart or MAQL and repaired lost context silently, and a stall was pushed past by the next simulated reply. A `context` mode adds, per turn: - The output must match the expected output: metrics, dimensions and every filter category for charts, MAQL for metrics, a subset of the arguments for tool calls. Alternatives are accepted; an end date after today counts as today; a single-value IN equals `=`. - A question back is graded by a binary judge (temperature 0, named boolean verdict): did it ask for information already in the conversation? Confirmation requests are not judged. - A reply that neither delivers nor asks is a stall; the user pushes once, then the turn ends. - The simulated user answers from the fixture's set answers and the conversation only, never from the expected output. Both modes compute every score; the mode (argument or GD_EVAL_CONVERSATION_MODE, default legacy) picks which verdict raises and is written as gate_passed. New scores: context_success, context_kept_rate, turns_before_first_break, lost_context_clarifications, stalled_turns. Fixtures may add depends_on, set_answers, expected_output_alternatives and expected_tool_args. Cleanup also deletes the alerts a conversation created, keeps metrics it only updated, and records objects from a stream that died mid-turn. fresh_conversation_per_turn runs a no-memory baseline for calibration. jira: QA-29448 risk: low
Unit tests for the paths of the context mode that had none: the judge's cache, an unreadable verdict, a provider fault announced once, the empty-body retry; the simulated user answering from facts, restating, and falling back without a key, on a provider error or an empty body; and the content check's edges (no chart or metric to compare, unreadable expected output or tool result, dimension and type mismatches). Drops a guard in _sort_signature the model already makes unreachable. jira: QA-29448 risk: nonprod
… patches Type annotations on the helpers and tests added for the context mode, as AGENTS.md asks of every function. The simulated-user tests patched `openai.OpenAI`, which fails before the code under test runs wherever the optional openai package is absent (a root `uv run pytest`); they now put a stand-in openai module in sys.modules, so they pass with and without the llm-judge extra. jira: QA-29448 risk: nonprod
test: improve multiturn conversation eval test
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## rel/dev #1837 +/- ##
===========================================
+ Coverage 83.19% 83.62% +0.43%
===========================================
Files 330 331 +1
Lines 21931 22338 +407
===========================================
+ Hits 18246 18681 +435
+ Misses 3685 3657 -28 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🚀 Automated PR to perform merge from master into rel/dev with changes up to ebe45f3 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/36830558397).