Skip to content

Repository files navigation

Agentic Engineering — A Curated Resource Guide

Opinionated. No tutorials, no listicles, no marketing. Continuously maintained — links verified weekly.


0. If you only read 7 things

  1. Building Effective Agents — Anthropic (Schluntz/Zhang). The "workflows vs agents" mental model that everything else builds on.
  2. Effective context engineering for AI agents — Anthropic. The successor concept to prompt engineering; defines the actual job.
  3. 12-Factor Agents — Dex Horthy / HumanLayer. Heroku's 12-factor reframed for LLM systems; the canonical "agents are mostly software" doctrine.
  4. How we built our multi-agent research system — Anthropic. The best single multi-agent case study, with concrete failure modes.
  5. Don't Build Multi-Agents — Cognition. Read alongside #4 — the productive disagreement at the heart of agent architecture in 2025/26.
  6. A practical guide to building agents (PDF) — OpenAI. The complementary canonical from the other lab.
  7. What We Learned From a Year of Building with LLMs — Yan/Bischof/Frye/Husain/Liu/Shankar. Tactical → operational → strategic; the field's distilled playbook.

1. Foundational design & "what is an agent"

2. Tool integration & MCP

3. Multi-agent orchestration frameworks

4. Durable execution for agents

5. Memory systems

6. Sandboxing & code execution

  • E2B docs — Firecracker microVM sandboxes; the de-facto hosted choice.
  • Modal Sandboxes — gVisor + filesystem snapshots; good for batch fleets.
  • Daytona docs — OSS sandbox repurposed for agents; sub-200ms cold start claims.
  • Cloudflare — Containers for Agents — Per-agent containers tied to Durable Objects.
  • cloudflare/sandbox-sdk — Reference SDK for spawning sandboxes from Workers.
  • apple/container — Native macOS container runtime; useful for local agent dev.
  • hyperlight-dev/hyperlight — Microsoft's sub-millisecond WASM/VM micro-sandbox.
  • gVisor docs — User-space kernel; understand it before trusting "sandboxed" claims.
  • Interpreters in Deep Agents: Code Between Tool Calls and Sandboxes — LangChain. Embedded interpreter runtimes let agents write code to coordinate tool calls, manage working state between steps, and control what gets surfaced into model context — reducing token pressure and enabling finer-grained orchestration than pure tool-dispatch.
  • smolmachines / smolvm as a sandbox for untrusted Python & JavaScript — Simon Willison. Tests a smolvm microVM as a sandbox for user-supplied Python/JS: RAM and CPU-time caps against `while true`, no network, and file access limited to chosen paths. Includes a reproducible test script. Finding: nested virtualization isn't available inside Firecracker-hosted agent containers, but GitHub Actions runners expose /dev/kvm, so the agent ran the tests there.

7. Inference & gateway infrastructure

8. Evaluation — philosophy (read these first)

9. Evaluation — frameworks & benchmarks

10. Observability & tracing

11. Production testing patterns & cost/latency

12. Security for agents

13. Coding agent infrastructure (read for harness design even if not building one)

14. SRE & operations agents (K8s, observability, IaC)


Worth following for ongoing signal

  • LangChain Blog — 7 Resources in the guide, mostly Coding agent infrastructure (read for harness design even if not building one)
  • Embrace The Red — 6 Resources in the guide, mostly Security for agents
  • Simon Willison's Weblog: coding-agents — 6 Resources in the guide, mostly Security for agents
  • Arize AI — 5 Resources in the guide, mostly Evaluation — frameworks & benchmarks
  • Elliot's Harness Lab | English — 4 Resources in the guide, mostly Security for agents
  • Eugene Yan — 4 Resources in the guide, mostly Evaluation — philosophy (read these first)
  • Hamel's Blog — 4 Resources in the guide, mostly Evaluation — philosophy (read these first)

About

A self-curating guide to agentic engineering.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages