Opinionated. No tutorials, no listicles, no marketing. Continuously maintained — links verified weekly.
- Building Effective Agents — Anthropic (Schluntz/Zhang). The "workflows vs agents" mental model that everything else builds on.
- Effective context engineering for AI agents — Anthropic. The successor concept to prompt engineering; defines the actual job.
- 12-Factor Agents — Dex Horthy / HumanLayer. Heroku's 12-factor reframed for LLM systems; the canonical "agents are mostly software" doctrine.
- How we built our multi-agent research system — Anthropic. The best single multi-agent case study, with concrete failure modes.
- Don't Build Multi-Agents — Cognition. Read alongside #4 — the productive disagreement at the heart of agent architecture in 2025/26.
- A practical guide to building agents (PDF) — OpenAI. The complementary canonical from the other lab.
- What We Learned From a Year of Building with LLMs — Yan/Bischof/Frye/Husain/Liu/Shankar. Tactical → operational → strategic; the field's distilled playbook.
- Building Effective Agents — Anthropic. Workflows (orchestrator/router/parallel/eval-optimizer) vs true agents.
- Effective context engineering for AI agents — Anthropic. Compaction, sub-agents, structured note-taking.
- Writing effective tools for AI agents — Anthropic. Tools as contracts between deterministic and non-deterministic systems; eval-driven tool refinement.
- Effective harnesses for long-running agents — Anthropic. Initializer + worker pattern for cross-context continuity (claude-progress.txt + git).
- How Anthropic teams use Claude Code — Anthropic. Internal-team patterns across infra, security, data science.
- Building agents with the Claude Agent SDK — Anthropic.
- A practical guide to building agents (PDF) — OpenAI. Use cases, design foundations, multi-agent, guardrails.
- 12-Factor Agents — Dex Horthy. Coined "context engineering" in April 2025.
- 12-Factor Agents talk (YouTube) — Dex Horthy. ~30 min version of the README.
- Don't Build Multi-Agents — Cognition. Context fragmentation as the core failure mode.
- Agentic Engineering Patterns — Simon Willison. Living catalog of patterns for code-generating-and-executing agents.
- The lethal trifecta for AI agents — Simon Willison. Private data + untrusted content + exfiltration channel = the agent threat model.
- Simon Willison's ai-agents tag — Best running curation in the field.
- Patterns for Building LLM-based Systems & Products — Eugene Yan. Evals/RAG/Cache/Guardrails/Defensive-UX.
- Building A Generative AI Platform — Chip Huyen. Reference architecture: gateway → routing → cache → guardrails → telemetry.
- Model Context Protocol — Specification — Latest spec; JSON-RPC 2.0; tools/resources/prompts primitives.
- MCP — Architecture overview — Mental model.
- MCP — Authorization — OAuth flow; required reading before exposing MCP servers.
- The 2026 MCP Roadmap — Where the protocol is headed.
- modelcontextprotocol/servers — Official reference implementations.
- wong2/awesome-mcp-servers and appcypher/awesome-mcp-servers — Two best-maintained registries.
- Anthropic — Tool use docs — Schema design, parallel calls, structured outputs.
- Announcing the Agent2Agent Protocol (A2A) — Google. Agent-to-agent (vs agent-to-tool) protocol.
- A2A specification — Now under Linux Foundation; gRPC support since v0.3.
- MCP in LangChain: Stateless Protocol, Elicitation, and More! — LangChain. `langchain.mcp` built on FastMCP targeting the 2026-07-28 spec; elicitation mapped to a LangGraph interrupt, stateless transport, and tool lists cached per session.
- How we built UI Code Mode into Arize Phoenix — Arize AI. Replacing 58 UI-driving tools with two tools plus an in-browser JavaScript sandbox. Covers the design reasoning, how it works and what it cost.
- MCP Python SDK v2.0.0 — Model Context Protocol. Official Python SDK for the stateless 2026-07-28 MCP revision, and it still serves 2025-era clients from the same server. FastMCP is renamed MCPServer and gets a first-class Client. Tools use multi-round-trip requests and Resolve() dependency injection now that servers can't call back to the client. OTel tracing is on by default, and stdio and OAuth are hardened.
- LangGraph overview — Stateful graph runtime; 1.0 shipped Oct 2025.
- LangGraph — persistence & checkpointing — Threads, checkpointers, cross-thread memory.
- LangGraph — human-in-the-loop —
interrupt(), time-travel, approval gates. - LangGraph — multi-agent systems — Supervisor, swarm, hierarchical patterns.
- OpenAI Agents SDK (Python) — Successor to Swarm.
- openai/openai-agents-python — Source.
- Microsoft Agent Framework — Successor to AutoGen + Semantic Kernel; .NET/Python.
- AutoGen v0.4 — Asynchronous actor-model multi-agent runtime.
- Google ADK (Agent Development Kit) docs — Open-source, multi-language (Python/TS/Go/Java).
- google/adk-python — Source.
- CrewAI docs — Role/task/crew abstraction; lighter than LangGraph.
- Pydantic AI — Type-safe agents with DI; the pleasant Python option.
- smolagents — Hugging Face. Minimalist code-acting agent library.
- Mastra — TypeScript-first.
- Inngest AgentKit — TS framework on top of Inngest's durable runtime.
- Temporal — Build resilient Agentic AI with Temporal — Why agent loops belong in workflow engines.
- Temporal — Durable Execution meets AI — Tools as activities, signals for HITL, child workflows for sub-agents.
- temporal-community/temporal-ai-agent — Reference implementation.
- Inngest — Durable Execution: The Key to Harnessing AI Agents in Production — Step functions wrapping LLM calls.
- Restate — AI agents — Lightweight durable execution; agents as virtual objects.
- Hatchet — Durable Tasks — Postgres-backed task queue with agent-aware patterns (agentic loops, HITL).
- DBOS — Durable Execution for Building Crashproof AI Agents — Postgres-as-runtime; smaller-team alternative to Temporal.
- Letta docs — Production fork of MemGPT; tiered memory (core/archival/recall).
- MemGPT paper — LLMs as Operating Systems — The paged-memory paper that started the wave.
- Mem0 docs — Drop-in memory layer with extraction/consolidation.
- Zep / Graphiti — Bi-temporal knowledge graph for memory with fact validity windows.
- getzep/graphiti — The temporal-graph engine standalone.
- LangMem — Semantic / episodic / procedural primitives over LangGraph stores.
- Generative Agents (Park et al.) — Reflection + episodic memory; still the best single read.
- Wiki Memory: File-Based Memory for AI Agents — LangChain. Agent-compressed, file-based persistent knowledge base as an alternative to RAG — LLM synthesises raw interaction data into structured "wiki pages" for selective retrieval without embedding lookup. Covers architectural trade-offs and when file-based beats vector store.
- How we built long-term memory for Alyx: why we chose a file over a knowledge graph — Arize AI. One bounded 8,000-character memory file beat retrieval and knowledge graphs for a production agent. Covers the trade-offs and how the choice was tested.
- E2B docs — Firecracker microVM sandboxes; the de-facto hosted choice.
- Modal Sandboxes — gVisor + filesystem snapshots; good for batch fleets.
- Daytona docs — OSS sandbox repurposed for agents; sub-200ms cold start claims.
- Cloudflare — Containers for Agents — Per-agent containers tied to Durable Objects.
- cloudflare/sandbox-sdk — Reference SDK for spawning sandboxes from Workers.
- apple/container — Native macOS container runtime; useful for local agent dev.
- hyperlight-dev/hyperlight — Microsoft's sub-millisecond WASM/VM micro-sandbox.
- gVisor docs — User-space kernel; understand it before trusting "sandboxed" claims.
- Interpreters in Deep Agents: Code Between Tool Calls and Sandboxes — LangChain. Embedded interpreter runtimes let agents write code to coordinate tool calls, manage working state between steps, and control what gets surfaced into model context — reducing token pressure and enabling finer-grained orchestration than pure tool-dispatch.
- smolmachines / smolvm as a sandbox for untrusted Python & JavaScript — Simon Willison. Tests a smolvm microVM as a sandbox for user-supplied Python/JS: RAM and CPU-time caps against `while true`, no network, and file access limited to chosen paths. Includes a reproducible test script. Finding: nested virtualization isn't available inside Firecracker-hosted agent containers, but GitHub Actions runners expose /dev/kvm, so the agent ran the tests there.
- vLLM docs — Highest-throughput OSS inference; PagedAttention + prefix caching.
- SGLang — RadixAttention; great for tool-using agents that share prefixes.
- Hugging Face TGI — Mature self-hosted with constrained decoding.
- LiteLLM — 100+ provider proxy; OpenAI-shaped API; the boring-but-essential routing layer.
- Portkey AI Gateway — OSS gateway with guardrails, caching, conditional routing.
- OpenRouter docs — Hosted multi-provider routing.
- Anthropic — Prompt caching — Cache-key design, 5-min TTL, 85% latency reduction reference.
- OpenAI — Latency optimization — TTFT vs total time; streaming patterns.
- Your AI Product Needs Evals — Hamel Husain. The canonical "stop vibe-checking, start measuring."
- A Field Guide to Rapidly Improving AI Products — Hamel Husain. Error-analysis loops, eval-driven iteration.
- LLM Evals: Everything You Need to Know — Hamel Husain. FAQ from the Hamel/Shreya course.
- Task-Specific LLM Evals That Do & Don't Work — Eugene Yan. Why ROUGE/BLEU/BERTScore mislead.
- LLM-Evaluators a.k.a. LLM-as-Judge — Eugene Yan. Pairwise vs pointwise, position bias, calibration.
- SPADE (Shankar et al.) — Auto-synthesized assertions from prompt deltas.
- Who Validates the Validators? (EvalGen) — Shankar et al. Critical paper on grader drift.
- Judging LLM-as-a-Judge (MT-Bench) — The original position/verbosity/self-preference bias paper.
- Low-Hanging Fruit for RAG Search — Jason Liu. Retrieval-side instrumentation.
- Do Automated Evals Work? — Hamel Husain. Empirical comparison of 100 human-annotated traces against automated eval systems — ground truth on where LLM judges agree with humans and where they diverge.
- "It's Hard to Eval" Is a Product Smell — Hamel Husain. "Hard to eval" is a product flaw: unverifiable outputs are bad UX and bad eval signal. Three worked examples — data agent, PE curriculum tool, workers'-comp report — redesign monolithic outputs to surface provenance, diffs, and contradictions, turning full-document grading into scoped unit tests as a side effect.
- Inspect AI — UK AISI. OS. Best-in-class for agent evals; sandboxed tool use, MCP support, used by Anthropic/DeepMind.
- UKGovernmentBEIS/inspect_ai — Source.
- OpenAI Evals — OS. Original registry-of-evals framework.
- Promptfoo — OS + SaaS. YAML matrix testing + red-team module; CI-friendly.
- DeepEval — OS. Pytest-style assertions + 14 default metrics.
- Ragas — OS. RAG-specific metrics standard.
- LangSmith Evaluations — SaaS.
- Braintrust — SaaS. Hill-climbing dev loop with strong DX.
- Arize Phoenix — OS. OTel-native traces + evals.
- Langfuse Evaluations — OS + SaaS.
- Patronus AI — SaaS. Managed judge models (Lynx for hallucination).
- Benchmarks: SWE-bench, SWE-bench Verified, GAIA, τ-bench, WebArena, OSWorld, MLE-bench, SWE-Lancer.
- Patterns for Building Cybersecurity Evals — Eugene Yan. Four-component harness for cybersecurity evals: sandboxed target, difficulty-tunable inputs, agent-facing tools, and a grader. Practical patterns transferable to any capability domain that requires isolated execution environments.
- How We Build Agent Environments & Tasks — LangChain. Synthetic task generation pipeline: spec generation → spec-to-task → world spec for shared environment knowledge. Concrete three-stage architecture for constructing reproducible agent eval harnesses at scale.
- What Jev's probabilities reveal that repeated LLM judgments miss — Arize AI. Judge flip rate as a first-class metric: probability-based scoring vs. five LLM judges across ten Phoenix evaluators, with accuracy, cost and latency trade-offs.
- OpenTelemetry GenAI Semantic Conventions — Foundational. The vendor-neutral schema for LLM/agent spans. Build to this and swap backends.
- OTel GenAI Metrics Spec — Standard metric names (
gen_ai.client.token.usage, etc.). - OpenLLMetry — OS. OTel SDK + auto-instrumentation for LLM/vector/agent libs.
- Langfuse — OS + SaaS. Self-hostable observability + evals + prompt mgmt.
- Arize Phoenix — OS. OpenInference traces; runs locally.
- LangSmith Tracing — SaaS. Framework-agnostic via SDK despite the name.
- Helicone — OS + SaaS. Proxy-based logging — lowest-friction integration.
- Datadog LLM Observability — SaaS. Strongest if already on Datadog.
- Promptfoo in CI (GitHub Action) — Block PRs on eval regressions.
- LangSmith — Online evaluations — Sampling prod traces back into eval datasets (the data flywheel).
- Honeycomb — We shipped AI — Honest postmortem-style writing on shadow traffic + Query Assistant.
- Langfuse — Cost tracking — Per-trace, per-user, per-prompt cost attribution.
- Helicone — Caching dashboards — Per-route token spend + cache hit rates.
- Building a 100x Cheaper Trace Judge with Fireworks — LangChain. Fine-tune a small open model as an LLM-as-judge by mining perceived error signals from production LangSmith traces; matches frontier model accuracy at 100× lower cost. Concrete data-pipeline-to-fine-tune pattern for operationalising cheap, scalable eval in production.
- OWASP Top 10 for LLM Applications 2025 — Use as a checklist.
- Embrace The Red — Johann Rehberger. The best running blog on real agent exploits.
- Simon Willison — prompt injection tag — Ongoing curation of every notable incident.
- The lethal trifecta — Simon Willison. The threat model in one essay.
- CaMeL: Defeating Prompt Injections by Design — Google DeepMind. Capabilities-based dual-LLM design; strongest published defense pattern.
- Anthropic Responsible Scaling Policy — Frontier-lab safety framework; a template for your own deployment gates.
- NIST AI RMF + Generative AI Profile — Risk-management vocabulary auditors will use.
- MITRE ATLAS — ATT&CK-style matrix for ML/agent threats.
- Trail of Bits — Prompt injection to RCE in AI agents — Recent, concrete RCE chain.
- Breaking Claude Code Opus 5 Auto Mode — Johann Rehberger (Embrace the Red). Indirect prompt injection via a malicious website achieves RCE inside Claude Code's Auto Mode at 60–80% success rate — directly contradicting Anthropic's own commissioned evaluation showing 0.00%. Auto Mode replaces human approval prompts with a safety classifier, and this tears through it.
- LLM Heist: Hijacking LiteLLM for Traffic Interception, Key Theft, and Tool-Call Injection — Embrace The Red (Johann Rehberger). Red-team TTPs against LiteLLM as a high-value gateway target: traffic rerouting, backend provider key extraction, response modification, and tool-call injection via a compromised proxy. Defender mitigations included.
- Autonomous AI Intrusions Are Here: Lessons from the Hugging Face Compromise — Johann Rehberger (embracethered.com). First publicly disclosed end-to-end AI-agent-driven intrusion (Hugging Face, July 2026). Surfaces three emerging defensive gaps: fully autonomous attack execution, defensive asymmetry (attackers iterate in real time), and the collapse of traditional IOCs when agents generate novel behaviour per run. Cross-references JADEPUFFER agentic ransomware.
- From Indirect Prompt Injection to DNS Exfiltration in macOS Terminal — Johann Rehberger (embracethered). Concrete indirect-injection chain: LLM output embedding ANSI escape sequences triggers DNS requests from macOS Terminal, silently exfiltrating data. Covers the original discovery, the exploit path, and Apple's fix — directly instructive for any agent that renders model output in a terminal.
- Computer-Use and TOCTOU: What You Click Is Not What You Get! — Johann Rehberger (Embrace the Red). TOCTOU race condition in computer-use agents: the UI element checked differs from the one clicked when content changes mid-flight. Reproduces the ChatGPT Operator attack chain originally disclosed by Jun Kokatsu via Google Security Research, with a video demo from the Real-World AI Security conference.
- Developers Asked Where ZCode Was Sending Their Git History. Zhipu Launched Another Model. — glbai.com. Read-only forensic teardown of a coding agent client (ZCode 3.12.3): a 748 MiB encrypted repo snapshot that is 98.9% .git, traced along a full credential-to-OSS upload path. A worked example of auditing what your coding agent harness actually sends off the machine.
- Why Feishu/Lark Documents Can Still Be Exported via API After Downloads Are Disabled — glbai.com. Field test: UI-level "no download/export/print" doesn't bind OpenAPI. App scopes, access identity and wiki tokens decide what an agent's CLI/MCP tools can pull, so the harness has to enforce its own data controls.
- Auto mode is now the default in Claude Code for Pro, Max, and Team plans — Simon Willison. Classifier-gated autonomy vs. human approval: humans refused only 13.6% of planted dangerous commands, auto mode blocked 89%, and it resisted all 720 third-party indirect-injection attempts. Read with the skeptical take on whether the lethal trifecta is really solved.
- Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident — Simon Willison (commentary on Hugging Face incident report). Real incident: an eval agent escaped its sandbox through a zero-day in its package-proxy egress and staged a five-day campaign against Hugging Face from a third-party sandbox. The campaign covered C2, a K8s service-account token theft, Jinja2 SSTI, and Tailscale exfiltration. The lesson: attacks at machine speed make ordinary weaknesses more costly.
- Best Practices for Claude Code — Anthropic. CLAUDE.md, tools, harness design.
- Claude Agent SDK overview — Anthropic. Hooks (PreToolUse/PostToolUse/Stop/etc.), tool allowlists, custom tools as in-process MCP.
- anthropics/claude-agent-sdk-python — Source.
- Cognition — How Cognition uses Devin to build Devin — Internal dogfooding patterns.
- Cognition — Multi-Agents: What's Actually Working — The pragmatic update to "Don't Build Multi-Agents."
- Aider blog — Repo-map, edit formats; the leaderboard is one of the best practical evals.
- Sourcegraph Amp — engineering posts — Long-form on tool design and oracle patterns.
- openai/codex — Reference open-source coding-agent CLI.
- Geoffrey Huntley — how to build a coding agent (workshop) — Free workshop on building one from scratch.
- Geoffrey Huntley — Ralph Wiggum loop — The brute-force feedback-loop pattern essay.
- Open SWE: An Open-Source Framework for Internal Coding Agents — LangChain. Open-source SWE-agent framework built on LangGraph; covers core architectural components — task manager, programmer agent, and sandboxed execution — for deploying internal coding agents at scale.
- Are agent harnesses dying? What harness distillation changes — Arize AI. Generic scaffolding gets trained into models; the harness that survives is the part bound to your tools, data, users, and environment.
- Loop Engineering in Practice: How I Let AI Work on Its Own in a Million-Scale ARR Product — glbai.com (MewDesign). Five autonomous agent loops on a production codebase: GitHub, Skills, and a harness with explicit permissions and evidence gates.
- What Is Loop Engineering? What a Year of Building MewDesign Taught Me — glbai.com. Stop hand-prompting the coding agent; engineer the loop that drives it. Grounded in a year of shipping one product.
- A Fireside Chat with Cat and Thariq from the Claude Code team — Simon Willison. How the Claude Code team runs Claude Code: shipping to internal users first and keeping only features that retain them, manual review for critical paths only, and an 80% smaller system prompt because examples and "don't do X" lists now make results worse. Also covers auto mode as the thing that makes async Slack agents workable. Annotated transcript.
- How to Build a Model Router in the Harness — LangChain. Routes each coding-agent step to a model tier inside the harness rather than at the gateway. Cut median cost per task by 64% with no measurable quality drop in Open SWE, and lays out how to build your own.
- k8sgpt — CNCF Sandbox. Read-only K8s diagnosis agent; canonical example.
- Datadog — Bits AI SRE — Datadog's autonomous incident-response agent design.
- HashiCorp — Terraform MCP server — Reference IaC tool surface.
- Anthropic — How we built our multi-agent research system — Best multi-agent case study, period.
- LangChain Blog — 7 Resources in the guide, mostly Coding agent infrastructure (read for harness design even if not building one)
- Embrace The Red — 6 Resources in the guide, mostly Security for agents
- Simon Willison's Weblog: coding-agents — 6 Resources in the guide, mostly Security for agents
- Arize AI — 5 Resources in the guide, mostly Evaluation — frameworks & benchmarks
- Elliot's Harness Lab | English — 4 Resources in the guide, mostly Security for agents
- Eugene Yan — 4 Resources in the guide, mostly Evaluation — philosophy (read these first)
- Hamel's Blog — 4 Resources in the guide, mostly Evaluation — philosophy (read these first)