Repository navigation
feat(sandbox): read_sandbox_file returns text pages and images - #3120
thomasbarrett wants to merge 2 commits into
Conversation
2ee6d65 to
ef175dc
Compare
1b544a6 to
bc6a734
Compare
bc6a734 to
ea71481
Compare
fe1f555 to
ab39df8
Compare
ab39df8 to
7fe6cd9
Compare
7fe6cd9 to
9d6078c
Compare
1329f1b to
2ea7213
Compare
read_sandbox_outputs and read_sandbox_file returned base64, which agents decoded by retyping it as output tokens. They now return text and images the model can read, and the Claude and Codex harnesses accept output lines large enough to carry those images. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Thomas Barrett <thomas@fluidstack.io>
2ea7213 to
abdcd2f
Compare
|
We can take this further and and other file-format specific tools like PDF or Jupyter notebook reading tools that the harness's native Read tools have so that they can work well in sandbox. Might also be worth tuning the sandbox read tool to more closely match the native read tool for each harness |
EItanya
left a comment
There was a problem hiding this comment.
🤖 AI-generated review.
Reviewed abdcd2f. Found two reproducible issues and one documentation improvement:
-
Output polling can corrupt valid UTF-8. sandboxes.go:139 skips index zero when checking incomplete characters. If
éarrives across two polls, the first poll consumes its leading byte; the combined output becomes��. Preserve incomplete characters during polling, including at index zero, and handle incomplete terminal output separately. -
Image encoding doesn’t enforce the 500 KB limit. sandboxes.go:264–273 returns the final JPEG even when every quality setting exceeds the budget. A valid 2000×2000 noise PNG produced 1,384,698 bytes, defeating the stated goal of avoiding downstream re-encoding. Reduce dimensions and retry, or explicitly report that the limit cannot be met. Add a detailed-image test that asserts encoded size.
-
Optional: correct the shared MCP instructions. prompts.go:18 still says all file transfers have a 1 MiB limit, potentially discouraging the larger reads this PR enables. Scope that statement to writes; the architecture docs also retain the obsolete base64 read contract.
MCP, Claude/Codex harness, and compiler tests passed. Both helper-level reproduction probes failed as described. Deployed E2E and live model behavior weren’t tested.
|
read_sandbox_outputs goes back to base64; only read_sandbox_file changes. PNG, JPEG and GIF images up to 10 MiB, the chat attachment limit, pass through unchanged, and Claude Code and Codex resize them for their models. Larger images ask the agent to write a smaller copy. The harnesses read CLI output lines up to 16 MiB instead of 1 MiB, so image results fit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Thomas Barrett <thomas@fluidstack.io>
read_sandbox_filereturns base64, which models can only read by decoding it as output tokens. It now behaves like Claude Code's nativeReadtool:N<tab>content), 2000 lines by default, paged withoffset/limit. A page also stops at 32 KiB, and lines over 2000 bytes are cut.The Claude and Codex harnesses now read output lines up to 16 MiB, up from 1 MiB, so image results fit.
Breaking:
read_sandbox_filereturns content instead ofdata_base64. It has only shipped in the 1.0.0 alphas.🤖 Generated with Claude Code