Skip to content

[bot] Merge master/ebe45f35 into rel/dev - #1837

Merged
yenkins-admin merged 6 commits into
rel/devfrom
snapshot-master-ebe45f35-to-rel/dev
Oct 1, 2026
Merged

yenkins-admin merged 6 commits into
rel/devfrom
snapshot-master-ebe45f35-to-rel/dev

Conversation

@yenkins-admin

Copy link
Copy Markdown
Contributor

🚀 Automated PR to perform merge from master into rel/dev with changes up to ebe45f3 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/36830558397).

myhoai and others added 6 commits October 1, 2026 09:33
gen-ai reports a turn that ended without a final answer as a 502 whose
error event carries a `reason` (max_iterations, max_tokens,
content_filter). The 502 alone routed it to the transient retry, which
re-sent the same question up to five times: the work was repeated, the
question landed in a conversation's history twice, and a stall was
hidden from every evaluator.

These reasons now raise TurnIncompleteError, a ChatError that is not
retried and keeps the partial result. A 502 without a reason, or with
`unknown`, is still treated as transient.

This changes behaviour for every eval kind, not only conversations: an
item whose turn hits the iteration limit now fails instead of being
retried.

jira: QA-29448
risk: low
create_metric is an upsert: on an existing id it replaces that metric
and reports `created_new: false`. Cleanup listed every create_metric
result as created, so a run in which the agent updated a metric that
was already in the workspace deleted it.

_extract_created_metric_ids now skips results with created_new false.

jira: QA-29448
risk: low
A conversation passed when every turn had the expected skill active and
produced an output of the expected type. Nothing compared what the
output contained (the check that tried returned None for every chart
in both conversation datasets), the simulated user was primed with the
expected chart or MAQL and repaired lost context silently, and a stall
was pushed past by the next simulated reply.

A `context` mode adds, per turn:

- The output must match the expected output: metrics, dimensions and
  every filter category for charts, MAQL for metrics, a subset of the
  arguments for tool calls. Alternatives are accepted; an end date
  after today counts as today; a single-value IN equals `=`.
- A question back is graded by a binary judge (temperature 0, named
  boolean verdict): did it ask for information already in the
  conversation? Confirmation requests are not judged.
- A reply that neither delivers nor asks is a stall; the user pushes
  once, then the turn ends.
- The simulated user answers from the fixture's set answers and the
  conversation only, never from the expected output.

Both modes compute every score; the mode (argument or
GD_EVAL_CONVERSATION_MODE, default legacy) picks which verdict raises
and is written as gate_passed. New scores: context_success,
context_kept_rate, turns_before_first_break,
lost_context_clarifications, stalled_turns. Fixtures may add
depends_on, set_answers, expected_output_alternatives and
expected_tool_args.

Cleanup also deletes the alerts a conversation created, keeps metrics
it only updated, and records objects from a stream that died mid-turn.
fresh_conversation_per_turn runs a no-memory baseline for calibration.

jira: QA-29448
risk: low
Unit tests for the paths of the context mode that had none: the judge's
cache, an unreadable verdict, a provider fault announced once, the
empty-body retry; the simulated user answering from facts, restating,
and falling back without a key, on a provider error or an empty body;
and the content check's edges (no chart or metric to compare,
unreadable expected output or tool result, dimension and type
mismatches). Drops a guard in _sort_signature the model already makes
unreachable.

jira: QA-29448
risk: nonprod
… patches

Type annotations on the helpers and tests added for the context mode,
as AGENTS.md asks of every function. The simulated-user tests patched
`openai.OpenAI`, which fails before the code under test runs wherever
the optional openai package is absent (a root `uv run pytest`); they now
put a stand-in openai module in sys.modules, so they pass with and
without the llm-judge extra.

jira: QA-29448
risk: nonprod
test: improve multiturn conversation eval test
@yenkins-admin
yenkins-admin merged commit 5288a90 into rel/dev Oct 1, 2026
1 check passed
@yenkins-admin
yenkins-admin deleted the snapshot-master-ebe45f35-to-rel/dev branch October 1, 2026 07:29
@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 68e02b4d-798b-4c13-b015-0c4b1c7216e5

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.78587% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 83.62%. Comparing base (6848b82) to head (ebe45f3).
⚠️ Report is 602 commits behind head on rel/dev.

Files with missing lines Patch % Lines
...val/src/gooddata_eval/core/agentic/conversation.py 99.69% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##           rel/dev    #1837      +/-   ##
===========================================
+ Coverage    83.19%   83.62%   +0.43%     
===========================================
  Files          330      331       +1     
  Lines        21931    22338     +407     
===========================================
+ Hits         18246    18681     +435     
+ Misses        3685     3657      -28     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants