Browsing: Field Notes

A practitioner framework for evaluating AI agents: a 21-point scorecard across 7 pillars (completion, accuracy, tool use, trajectory, reliability, latency/cost, safety), a three-test loop you can run in an afternoon, an interactive calculator, and a comparison of every major eval tool in 2026. Built for U.S. teams shipping agents at small and mid-sized companies.

MCP, the Model Context Protocol, is the open standard that lets AI apps like Claude Desktop, Cursor, and Windsurf connect to your files, your GitHub, your Notion, and any other tool – without a custom integration per app. A plain-English beginner guide with an analogy-first walkthrough, a worked example, a comparison against plugins and function calling, and a no-hype FAQ.