Repository navigation
Conversation
Add demonstration-based setup for computer-use agents
Deploy the demonstration setup UI and recorder images
Generate the UI client in clean builds
* fix(setup): handle null recorder step fields * fix(setup): collect downloaded files * fix(setup): block credential URL queries * fix(setup): reject URLs with query data * fix(recording): reject URLs with query data * fix(setup): validate navigation URL fallbacks * fix(setup): never retain demonstrated values * fix(setup): reject credential intent * fix(setup): guard all credential boundaries * fix(setup): disable runtime form arguments * docs(setup): describe form-entry restriction * docs(setup): define empty argument contract * fix(setup): type empty host arguments --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(setup): match credential intent by term * fix(setup): block MFA and hyphenated sign-in intent * fix(setup): normalize credential intent matching * fix(setup): normalize credential intent paths * fix(setup): reject conventional PIN values * fix(setup): normalize separated credential terms * fix(setup): normalize credential intent tokens * fix(setup): refine credential intent grammar * fix(setup): block concatenated sign-in intent * fix(setup): normalize credential separators * fix(setup): scope credential intent matching * fix(setup): preserve credential guard boundaries * fix(setup): normalize credential path separators * fix(setup): tokenize credential intent consistently * fix(setup): validate raw credential intent * fix(setup): constrain credential exceptions * fix(setup): scope recorder hostname exception * fix(setup): reject PIN and backup codes * fix(setup): distinguish PIN codes from pin actions * fix(setup): narrow PIN action exemption * fix(setup): constrain pin action destinations * fix(setup): require exact pin action destination * fix(setup): scope pin action exception * fix(setup): close remaining pin intent gaps * fix(setup): separate pin credentials from content * fix(setup): scope safe pin identifiers * fix(setup): reject bare PIN credential actions * fix(setup): narrow PIN context exceptions * fix(setup): close credential boundary bypasses * test(setup): cover credential boundary regressions * fix(setup): reject numeric credential suffixes * test(setup): cover prefixed credential values * fix(setup): reject prefixed credential values * test(setup): cover concatenated credential prefixes * fix(setup): cover concatenated credential values * test(setup): cover credential prefix heuristics * fix(setup): refine credential token detection * test(setup): cover credential token boundaries * fix(setup): bound concatenated credentials * test(ui): cover structural credential boundaries * fix(ui): classify structural credential intent * test(ui): cover mixed credential structures * fix(ui): classify mixed credential structures --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* Record sign-in fields as credential placeholders Sign-in fields become credential steps that carry only their kind (username, password, otp). The compiler turns them into exact $placeholders, and the host collects and stores the values, so the setup page, recorder and model never receive them. * fix(recording): keep sign-in submit clicks; let the host decide re-prompts Review: a submit button inside a password form was classified as a username field, so the Sign in click was dropped. Only typed fields are credential fields now. The setup page always asks the host, which binds saved sign-in details to the demonstrated website. * fix(setup): reject literal sign-in values; ignore checkbox input events Review: removing the word heuristic let a goal like 'password example-secret-123' reach the prompt. A narrow check now rejects a credential word followed by an assigned or digit-bearing value, while sign-in wording and $placeholders stay allowed. Checkbox and radio input events no longer become form-entry steps. * fix(setup): classify secret fields before the username fallback; keep reviewed descriptions Review: a 'Secret key' field in a password form became $username, and an edited credential step description was dropped from the prompt. * fix(setup): reject symbol-bearing sign-in values and credential URL paths Review: 'password correct-horse-battery-staple', 'verification code 482913' and /token/<value> paths compiled into prompts. * fix(setup): check raw and repeatedly decoded URL paths; labelled codes Review: /token/abc123/../reports and %2574oken paths, and 'code sent to me: 482913', bypassed the disclosure check. * fix(setup): treat backslashes as path separators when checking URLs Review: https://host\token\abc123\..\reports resolved to a clean path while the saved URL kept the token segment. * fix(setup): classify password and username fields by type first; nearby codes Review: a password field labelled Passcode became an OTP step, and an autocomplete=username field with 'auth' in its id became a password. Credential steps can now be reclassified in review. Codes a few words after their label are rejected. --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
Each step is one row: number, instruction, optional expected outcome and an icon remove button. A ten-step demonstration now fits on a 1440x900 screen with the actions visible; narrow screens stack the outcome under its instruction. Co-authored-by: Joonatan Samuel <jonksar@github.com>
Co-authored-by: Joonatan Samuel <jonksar@github.com>
* Support form entry in demonstrations Typed text and dropdown choices were dropped at capture, so the agent could not repeat them and setup rejected every form step. Ordinary fields now keep their text and selects keep the chosen label; the compiler turns them into exact 'replace with' / 'choose' instructions editable in review. Sign-in fields still record only their kind, with a server-side label backstop. Switching a field back from a sign-in kind resets its description. * fix(recording): keep typed text verbatim; treat PIN and passphrase fields as secrets Review: PIN and passphrase fields kept their values, and typed text was whitespace-normalized and truncated before being replayed 'exactly'. Text over 2000 characters is now flagged for the user instead of cut. The out-of-date host message now says to reload Reiterate. * fix(setup): catch revealed password fields; keep multiline values; sync edited choices Review: a password field revealed as type=text kept its value, the single-line review input flattened multiline text, and editing a choice left the recorded 'Choose PDF' description contradicting the new option. * fix(setup): replay multi-select choices as separate options Review: joining selected labels with commas produced an option that does not exist. Multi-select values are a JSON array of labels and compile to 'select exactly these options and no others'. * style(recording): wrap long multi-select line --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
- Show 'Changes require a new test.' only when a test result is discarded, not on the first keystroke of Describe. - Assume https:// for a website address typed without a scheme. - Pick the daily run time in the user's local time zone; the host still receives a UTC cron, and the UTC time is shown under the field. - Show a failed test as an error with a next step, and label the button 'Run test again'. - Disable the run button while a test is running. - Give headings a heading weight; make 'Change sign-in details' a normal link instead of a destructive red. Co-authored-by: Joonatan Samuel <jonksar@github.com>
…ion with scrolling steps and clipboard (#18) Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): add web-agent test feedback and completion checks * fix(setup): preserve schedule retries and match text stage contract * fix(setup): suppress ambiguous email route replacement * review round2 fixes * fix(setup): guard export email rewrites and unchanged custom blur * fix(setup): don't claim no steps ran when the AI service stops mid-run * fix(setup): keep recipient checks and custom done-when reselectable --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(web-agent-edit): add existing agent editor * fix(web-agent-edit): editor layout and review fixes * test(web-agent-edit): control load recovery fixture --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): allow free-text success criteria for Done when * fix(setup): keep described evidence out of exact-text suggestions --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(ui): compact agent setup header and describe step * fix(ui): keep stepper and close on one header row at narrow widths --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
…oter (#24) Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(ui): clarify the web agent edit page * fix(ui): keep edit dialogs and running tests guarded * fix(ui): keep conflict actions on one row and let maximum actions be typed --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(ui): readable agent screen bar and compact test heading - Parse each screen's thought into a typed AgentThought (final outcome, reasoning summaries, proposed actions, replay, text) instead of printing the raw string, so the outcome JSON is never shown. - AgentScreenBar shows a kind badge, one truncated line, More/Less for the rest (confirmation, other reasoning, safety checks) and icon paging. - Test stage heading is now "Verify agent can follow the process" with the description behind a (?) HelpTip. * test(ui): keep the agent-notes mock on its own line * docs(setup): note the test-stage HelpTip and how screen thoughts are read * fix(ui): polish the test stage to the design floor - Screen bar: only paint the agent's "completed" green when the host passed the test; announce paging politely; plain action text instead of chip boxes; label the pager group. - HelpTip: drawn icon instead of a "?" glyph, 28 px target, short shadow, z-index 1 under the setup dialogs, bubble capped to the viewport. - Test rail: drop the 3 px red side stripes on the failed step and failed Done-when (tint + red marker carry it) and the diagonal stripe empty browser. e2e asserts no side stripes in either failure. --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
* fix(ui): show edit test evidence and precise raw changes * fix(ui): plain failure labels and change counts on the edit page --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
…led (#29) Co-authored-by: Joonatan Samuel <jonksar@github.com>
* feat(setup): group demonstrated steps into stages and show downloads After a demonstration stops, the recorder asks gpt-5.6-luna to group the steps into short stages (Sign in, Open reports, Download the report) and reword each step. The stop response returns at once; the setup page polls until organizing finishes and keeps any edits made meanwhile. Clicks now name the control rather than its container, so a click inside a sign-in panel no longer records the whole panel's text. A double click is one step. The live view has no download bar, so the recorder reports each browser download (started, completed, failed). The setup page shows it over the browser, and the download becomes a step the agent confirms instead of repeating. * fix(setup): address review of stages and downloads - Name clicks inside web components (composedPath), table rows and focusable elements; a label click is recorded once, as its field. - Count quick repeat clicks on one control as one step with a count instead of dropping them, so paging a calendar replays correctly. - Tell the agent the download is kept and reported in a message, not on screen, and to not start it again. Drop failed downloads. - Keep the recorded label next to a reworded click. - Do not send typed text or chosen options to the organizer. - Rename stages on the edit page; a moved step joins the stage it moves into, and later unstaged steps get a neutral heading. --------- Co-authored-by: Joonatan Samuel <jonksar@github.com>
|
Opened against the wrong repository by mistake, sorry. |
|
| GitGuardian id | GitGuardian status | Secret | Commit | Filename | |
|---|---|---|---|---|---|
| - | - | Generic Password | e04e326 | recording/tests/test_api.py | View secret |
| - | - | Username Password | 64cbfd0 | ui/src/setup/AgentSetup.tsx | View secret |
🛠 Guidelines to remediate hardcoded secrets
- Understand the implications of revoking this secret by investigating where it is used in your code.
- Replace and store your secrets safely. Learn here the best practices.
- Revoke and rotate these secrets.
- If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.
To avoid such incidents in the future consider
- following these best practices for managing and storing secrets including API keys and other credentials
- install secret detection on pre-commit to catch secret before it leaves your machine and ease remediation.
🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.
Brings the edit page's top bar in line with the editor's workflow top bar (TopNavBar): one 64px row, 18px/500 title.
Before/after at 1440 and 900 px checked; full e2e suite 119/119.
Summary by cubic
Adds a demonstration-based setup flow for computer-use agents and replaces the default app: users describe a task, demonstrate it in a remote browser, review the captured steps, test, and schedule. Previously the
ui/app rendered the ReactFlow workflow editor; now it renders an embedded setup wizard, with a newrecording/capture service and production images beside it.New Features
recording/captures demonstrations in ephemeral Browserbase sessions, groups steps into stages and rewording via an LLM, reports downloads, and stores credential fields as kind-only placeholders so values never reach the UI or model.Dependencies
pytestfor the recorder; pushes to main build and pushuiandrecordingimages to ECR.docs/agent-setup-host.md; the app requires aparentOriginquery parameter.Written for commit f29096d. Summary will update on new commits.