Skip to content

Workflow-style top bar on the agent edit page - #184

Closed
Jonksar wants to merge 28 commits into
browser-use:mainfrom
iter8-ai:hermes/edit-topbar
Closed

Jonksar wants to merge 28 commits into
browser-use:mainfrom
iter8-ai:hermes/edit-topbar

Conversation

@Jonksar

@Jonksar Jonksar commented Oct 3, 2026 •

Copy link
Copy Markdown

Brings the edit page's top bar in line with the editor's workflow top bar (TopNavBar): one 64px row, 18px/500 title.

  • Agent name is an inline-editable title (borderless until hover/focus) instead of a labelled input box; "Agent name" and "Name saves automatically" stay as visually-hidden label/description, so a11y and e2e selectors are unchanged.
  • Version is a rounded chip, schedule sits after a thin divider and ellipsises instead of wrapping.
  • Rename status (Saving…/Saved) and rename errors unchanged.

Before/after at 1440 and 900 px checked; full e2e suite 119/119.


Summary by cubic

Adds a demonstration-based setup flow for computer-use agents and replaces the default app: users describe a task, demonstrate it in a remote browser, review the captured steps, test, and schedule. Previously the ui/ app rendered the ReactFlow workflow editor; now it renders an embedded setup wizard, with a new recording/ capture service and production images beside it.

New Features

  • Setup wizard covers describe, demonstrate, review, test, and schedule; the host owns auth, storage, execution, and scheduling through an origin-checked message bridge.
  • recording/ captures demonstrations in ephemeral Browserbase sessions, groups steps into stages and rewording via an LLM, reports downloads, and stores credential fields as kind-only placeholders so values never reach the UI or model.
  • Recording accepts only public HTTP(S) URLs with no query parameters or fragments and no embedded credentials.
  • The edit page's top bar matches the editor's workflow bar: inline-editable agent title, version chip, and ellipsizing schedule after a divider; rename status and errors are unchanged, and visually-hidden labels keep a11y selectors stable.

Dependencies

  • CI now runs UI unit tests, builds, and Playwright e2e plus pytest for the recorder; pushes to main build and push ui and recording images to ECR.
  • The host contract for embedding the wizard is documented in docs/agent-setup-host.md; the app requires a parentOrigin query parameter.

Written for commit f29096d. Summary will update on new commits.

Review in cubic

Joonatan Samuel and others added 28 commits September 11, 2026 12:24
Add demonstration-based setup for computer-use agents
Deploy the demonstration setup UI and recorder images
Generate the UI client in clean builds
* fix(setup): handle null recorder step fields

* fix(setup): collect downloaded files

* fix(setup): block credential URL queries

* fix(setup): reject URLs with query data

* fix(recording): reject URLs with query data

* fix(setup): validate navigation URL fallbacks

* fix(setup): never retain demonstrated values

* fix(setup): reject credential intent

* fix(setup): guard all credential boundaries

* fix(setup): disable runtime form arguments

* docs(setup): describe form-entry restriction

* docs(setup): define empty argument contract

* fix(setup): type empty host arguments

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(setup): match credential intent by term

* fix(setup): block MFA and hyphenated sign-in intent

* fix(setup): normalize credential intent matching

* fix(setup): normalize credential intent paths

* fix(setup): reject conventional PIN values

* fix(setup): normalize separated credential terms

* fix(setup): normalize credential intent tokens

* fix(setup): refine credential intent grammar

* fix(setup): block concatenated sign-in intent

* fix(setup): normalize credential separators

* fix(setup): scope credential intent matching

* fix(setup): preserve credential guard boundaries

* fix(setup): normalize credential path separators

* fix(setup): tokenize credential intent consistently

* fix(setup): validate raw credential intent

* fix(setup): constrain credential exceptions

* fix(setup): scope recorder hostname exception

* fix(setup): reject PIN and backup codes

* fix(setup): distinguish PIN codes from pin actions

* fix(setup): narrow PIN action exemption

* fix(setup): constrain pin action destinations

* fix(setup): require exact pin action destination

* fix(setup): scope pin action exception

* fix(setup): close remaining pin intent gaps

* fix(setup): separate pin credentials from content

* fix(setup): scope safe pin identifiers

* fix(setup): reject bare PIN credential actions

* fix(setup): narrow PIN context exceptions

* fix(setup): close credential boundary bypasses

* test(setup): cover credential boundary regressions

* fix(setup): reject numeric credential suffixes

* test(setup): cover prefixed credential values

* fix(setup): reject prefixed credential values

* test(setup): cover concatenated credential prefixes

* fix(setup): cover concatenated credential values

* test(setup): cover credential prefix heuristics

* fix(setup): refine credential token detection

* test(setup): cover credential token boundaries

* fix(setup): bound concatenated credentials

* test(ui): cover structural credential boundaries

* fix(ui): classify structural credential intent

* test(ui): cover mixed credential structures

* fix(ui): classify mixed credential structures

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* Record sign-in fields as credential placeholders

Sign-in fields become credential steps that carry only their kind
(username, password, otp). The compiler turns them into exact $placeholders,
and the host collects and stores the values, so the setup page, recorder and
model never receive them.

* fix(recording): keep sign-in submit clicks; let the host decide re-prompts

Review: a submit button inside a password form was classified as a
username field, so the Sign in click was dropped. Only typed fields are
credential fields now. The setup page always asks the host, which binds
saved sign-in details to the demonstrated website.

* fix(setup): reject literal sign-in values; ignore checkbox input events

Review: removing the word heuristic let a goal like 'password
example-secret-123' reach the prompt. A narrow check now rejects a
credential word followed by an assigned or digit-bearing value, while
sign-in wording and $placeholders stay allowed. Checkbox and radio
input events no longer become form-entry steps.

* fix(setup): classify secret fields before the username fallback; keep reviewed descriptions

Review: a 'Secret key' field in a password form became $username, and an
edited credential step description was dropped from the prompt.

* fix(setup): reject symbol-bearing sign-in values and credential URL paths

Review: 'password correct-horse-battery-staple', 'verification code
482913' and /token/<value> paths compiled into prompts.

* fix(setup): check raw and repeatedly decoded URL paths; labelled codes

Review: /token/abc123/../reports and %2574oken paths, and 'code sent to
me: 482913', bypassed the disclosure check.

* fix(setup): treat backslashes as path separators when checking URLs

Review: https://host\token\abc123\..\reports resolved to a clean path
while the saved URL kept the token segment.

* fix(setup): classify password and username fields by type first; nearby codes

Review: a password field labelled Passcode became an OTP step, and an
autocomplete=username field with 'auth' in its id became a password.
Credential steps can now be reclassified in review. Codes a few words
after their label are rejected.

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
Each step is one row: number, instruction, optional expected outcome and
an icon remove button. A ten-step demonstration now fits on a 1440x900
screen with the actions visible; narrow screens stack the outcome under
its instruction.

Co-authored-by: Joonatan Samuel <jonksar@github.com>
Co-authored-by: Joonatan Samuel <jonksar@github.com>
* Support form entry in demonstrations

Typed text and dropdown choices were dropped at capture, so the agent
could not repeat them and setup rejected every form step. Ordinary fields
now keep their text and selects keep the chosen label; the compiler turns
them into exact 'replace with' / 'choose' instructions editable in review.
Sign-in fields still record only their kind, with a server-side label
backstop. Switching a field back from a sign-in kind resets its
description.

* fix(recording): keep typed text verbatim; treat PIN and passphrase fields as secrets

Review: PIN and passphrase fields kept their values, and typed text was
whitespace-normalized and truncated before being replayed 'exactly'.
Text over 2000 characters is now flagged for the user instead of cut.
The out-of-date host message now says to reload Reiterate.

* fix(setup): catch revealed password fields; keep multiline values; sync edited choices

Review: a password field revealed as type=text kept its value, the
single-line review input flattened multiline text, and editing a choice
left the recorded 'Choose PDF' description contradicting the new option.

* fix(setup): replay multi-select choices as separate options

Review: joining selected labels with commas produced an option that does
not exist. Multi-select values are a JSON array of labels and compile to
'select exactly these options and no others'.

* style(recording): wrap long multi-select line

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
- Show 'Changes require a new test.' only when a test result is discarded,
  not on the first keystroke of Describe.
- Assume https:// for a website address typed without a scheme.
- Pick the daily run time in the user's local time zone; the host still
  receives a UTC cron, and the UTC time is shown under the field.
- Show a failed test as an error with a next step, and label the button
  'Run test again'.
- Disable the run button while a test is running.
- Give headings a heading weight; make 'Change sign-in details' a normal
  link instead of a destructive red.

Co-authored-by: Joonatan Samuel <jonksar@github.com>
…ion with scrolling steps and clipboard (#18)

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): add web-agent test feedback and completion checks

* fix(setup): preserve schedule retries and match text stage contract

* fix(setup): suppress ambiguous email route replacement

* review round2 fixes

* fix(setup): guard export email rewrites and unchanged custom blur

* fix(setup): don't claim no steps ran when the AI service stops mid-run

* fix(setup): keep recipient checks and custom done-when reselectable

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(web-agent-edit): add existing agent editor

* fix(web-agent-edit): editor layout and review fixes

* test(web-agent-edit): control load recovery fixture

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): allow free-text success criteria for Done when

* fix(setup): keep described evidence out of exact-text suggestions

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(ui): compact agent setup header and describe step

* fix(ui): keep stepper and close on one header row at narrow widths

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
…oter (#24)

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(ui): clarify the web agent edit page

* fix(ui): keep edit dialogs and running tests guarded

* fix(ui): keep conflict actions on one row and let maximum actions be typed

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(ui): readable agent screen bar and compact test heading

- Parse each screen's thought into a typed AgentThought (final outcome,
  reasoning summaries, proposed actions, replay, text) instead of printing
  the raw string, so the outcome JSON is never shown.
- AgentScreenBar shows a kind badge, one truncated line, More/Less for the
  rest (confirmation, other reasoning, safety checks) and icon paging.
- Test stage heading is now "Verify agent can follow the process" with the
  description behind a (?) HelpTip.

* test(ui): keep the agent-notes mock on its own line

* docs(setup): note the test-stage HelpTip and how screen thoughts are read

* fix(ui): polish the test stage to the design floor

- Screen bar: only paint the agent's "completed" green when the host
  passed the test; announce paging politely; plain action text instead of
  chip boxes; label the pager group.
- HelpTip: drawn icon instead of a "?" glyph, 28 px target, short shadow,
  z-index 1 under the setup dialogs, bubble capped to the viewport.
- Test rail: drop the 3 px red side stripes on the failed step and failed
  Done-when (tint + red marker carry it) and the diagonal stripe empty
  browser. e2e asserts no side stripes in either failure.

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(ui): show edit test evidence and precise raw changes

* fix(ui): plain failure labels and change counts on the edit page

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
)

* feat(ui): make agent instructions the main edit area with sign-in details beside it

* fix(ui): keep instructions beside the rail on smaller windows

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
…led (#29)

Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): group demonstrated steps into stages and show downloads

After a demonstration stops, the recorder asks gpt-5.6-luna to group the
steps into short stages (Sign in, Open reports, Download the report) and
reword each step. The stop response returns at once; the setup page polls
until organizing finishes and keeps any edits made meanwhile.

Clicks now name the control rather than its container, so a click inside
a sign-in panel no longer records the whole panel's text. A double click
is one step.

The live view has no download bar, so the recorder reports each browser
download (started, completed, failed). The setup page shows it over the
browser, and the download becomes a step the agent confirms instead of
repeating.

* fix(setup): address review of stages and downloads

- Name clicks inside web components (composedPath), table rows and
  focusable elements; a label click is recorded once, as its field.
- Count quick repeat clicks on one control as one step with a count
  instead of dropping them, so paging a calendar replays correctly.
- Tell the agent the download is kept and reported in a message, not
  on screen, and to not start it again. Drop failed downloads.
- Keep the recorded label next to a reworded click.
- Do not send typed text or chosen options to the organizer.
- Rename stages on the edit page; a moved step joins the stage it
  moves into, and later unstaged steps get a neutral heading.

---------

Co-authored-by: Joonatan Samuel <jonksar@github.com>
@Jonksar

Jonksar commented Oct 3, 2026

Copy link
Copy Markdown
Author

Opened against the wrong repository by mistake, sorry.

@Jonksar Jonksar closed this Oct 3, 2026
@gitguardian

gitguardian Bot commented Oct 3, 2026

Copy link
Copy Markdown

⚠️ GitGuardian has uncovered 2 secrets following the scan of your pull request.

Please consider investigating the findings and remediating the incidents. Failure to do so may lead to compromising the associated services or software components.

Since your pull request originates from a forked repository, GitGuardian is not able to associate the secrets uncovered with secret incidents on your GitGuardian dashboard.
Skipping this check run and merging your pull request will create secret incidents on your GitGuardian dashboard.

🔎 Detected hardcoded secrets in your pull request
GitGuardian id GitGuardian status Secret Commit Filename
- - Generic Password e04e326 recording/tests/test_api.py View secret
- - Username Password 64cbfd0 ui/src/setup/AgentSetup.tsx View secret
🛠 Guidelines to remediate hardcoded secrets
  1. Understand the implications of revoking this secret by investigating where it is used in your code.
  2. Replace and store your secrets safely. Learn here the best practices.
  3. Revoke and rotate these secrets.
  4. If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.

To avoid such incidents in the future consider


🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant