
Pick an AI agent framework by answering 3 questions: do your agents need durable state that survives a crash or a human pause, will you run models from more than 1 vendor, and how much API churn can your team absorb. If you need durable state, the shortlist is the graph runtimes that have reached 1.0: LangGraph, Google's Agent Development Kit and Microsoft Agent Framework. If you want a thin, typed agent loop, it is Pydantic AI, the OpenAI Agents SDK or smolagents. If you are committed to Claude, the Claude Agent SDK gives you Claude Code's own loop, and supports nothing else. Everything else is detail hanging off those 3.
Most comparisons start in the wrong place, with GitHub stars and a 20 line demo. The demo always works. What differs is what the framework does to your project for the next 2 years. We installed 10 Python frameworks on 6 October 2026: installed size ranged from 34 MB to 677 MB, and stable releases in the past 12 months from 4 to 192. The failure is ordinary: on an Intel Mac, pip install crewai installs a version 8 months old without a warning, and the team debugs an API the documentation no longer describes.
Durable state is the question that splits the field. An agent that runs for 5 seconds and answers a question needs almost nothing from a framework. An agent that runs for an hour, waits for a human to approve a refund, and must resume after a deploy needs checkpointing, resumable runs and an execution model that knows where it stopped. LangGraph, Google ADK 2 and Microsoft Agent Framework build that in as a graph of steps. Pydantic AI and the OpenAI Agents SDK get it from integrations with workflow engines such as Temporal or DBOS.
Model choice is the question teams answer too late. 9 of the 10 frameworks we checked document support for models from several vendors. The Claude Agent SDK documents Claude only, across the Anthropic API and the clouds that host Claude. The OpenAI Agents SDK supports other providers through OpenAI-compatible endpoints and adapters, 2 of which its documentation marks as beta. If you expect to switch models when prices or quality change, that matters on day 1.
Churn is the question nobody measures. A framework that ships a release every 2 days is either fixing things quickly or moving the ground under you, and the version number tells you which. 3 of the 10 are still below 1.0: the OpenAI Agents SDK, the Claude Agent SDK and AutoGen.

Answer the 3, then spike 2 candidates. Build the same small agent on both, with 1 tool, 1 human approval and 1 failure, and keep the one your team could debug at 2 in the morning.
We installed each framework into its own clean Python 3.12 virtual environment on 6 October 2026, on a 2019 MacBook Pro with an Intel i7-9750H, and recorded what pip brought with it. Import time is the median of 5 runs of each framework's main entry point after the first run cached its bytecode.
Framework | Version | 1.0 since | Licence | Other packages | Installed MB | Import, seconds |
LangGraph | 1.2.13 | Oct 2025 | MIT | 38 | 59 | 0.69 |
Google ADK | 2.11.0 | May 2025 | Apache-2.0 | 48 | 104 | 1.24 |
Microsoft Agent Framework | 1.20.0 | Apr 2026 | MIT | 206 (core: 11) | 677 (core: 34) | 0.18 (core: 0.20) |
Pydantic AI | 2.54.0 | Sep 2025 | MIT | 97 (slim: 17) | 209 (slim: 47) | 1.94 (slim: 0.76) |
OpenAI Agents SDK | 0.23.1 | Not yet | MIT | 38 | 90 | 2.05 |
smolagents | 1.26.0 | Dec 2024 | Apache-2.0 | 28 | 75 | 0.43 |
Agno | 3.1.1 | Jan 2025 | Apache-2.0 | 31 | 98 | 0.39 |
CrewAI | 1.9.3 installed, 1.15.23 latest | Oct 2025 | MIT | 131 | 555 | 1.68 |
Claude Agent SDK | 0.2.163 | Not yet | MIT | 30 | 278 | 0.93 |
AutoGen AgentChat | 0.7.5 | Not yet | MIT | 11 | 44 | 0.29 |

The default package is often not the one you want. agent-framework installs Microsoft's framework with every integration: 206 other packages and 677 MB. agent-framework-core installs the same Agent class with 11 other packages and 34 MB. Pydantic AI does the same thing: the full package pulls in 97 packages and 209 MB, the pydantic-ai-slim package 17 and 47 MB. Install the core package and add only the integrations you call.
The Claude Agent SDK's size is a binary, not a dependency tree. Of its 278 MB, 223 MB is a single bundled claude executable: the SDK runs Claude Code's own agent loop as a subprocess. That is the point of the SDK, and it is also why it behaves more like a runtime than a library.
CrewAI's 555 MB comes from its dependencies, not its own code. The largest directories in its environment were the Kubernetes client, ONNX Runtime, SymPy and the Chroma vector database bindings. You install them whether or not you use the features that need them.
Import times are small and mostly irrelevant, with 1 exception. A server imports once. A command-line tool or a serverless function imports on every cold start, and there the gap between 0.18 seconds and 2.05 seconds on this machine is worth knowing.
We counted every version each framework published on PyPI between 6 October 2025 and 6 October 2026, excluding yanked releases and splitting stable versions from pre-releases and nightly builds.

Pydantic AI shipped 192 stable releases in 12 months, roughly 1 every 2 days, and still moved to a 2.0 major version in June 2026. Agno shipped 128 stable releases and reached 3.0 in August 2026. Both are fast-moving projects that version honestly, which means your lock file matters more than your requirements file.
The 2 vendor SDKs below 1.0 release almost as often. The Claude Agent SDK published 108 stable versions, with 38 more versions yanked entirely, and the OpenAI Agents SDK 87. Below 1.0, semantic versioning promises nothing about compatibility between minor releases, so every upgrade is a small migration you should test.
CrewAI's 290 is mostly nightly builds. 59 were stable and 231 were pre-release or nightly versions, most of them dated nightly builds. Nightly builds are not a sign of instability on their own, but they are another reason to pin exact versions.
AutoGen published nothing, because it has stopped. Microsoft's AutoGen repository states that the project is in maintenance mode, will receive no new features, and directs new users to Microsoft Agent Framework, which reached 1.0 on 2 April 2026 and describes itself as the direct successor to both AutoGen and Semantic Kernel. In the Stack Overflow 2025 Developer Survey, 12 per cent of developers who used agent frameworks had used AutoGen. Those projects now have a migration ahead of them.
smolagents is the opposite end. 4 stable releases in a year, a deliberately small core, and the slowest-moving API in the set.
These 3 treat an agent as a graph of steps with explicit state, and that is what makes long-running work recoverable. Each documents built-in checkpointing or resumption, human-in-the-loop interrupts, tracing and support for the Model Context Protocol.
LangGraph was first to a stable release in this category. Its 1.0 shipped on PyPI on 17 October 2025, and LangChain's announcement described it as the first stable major release among durable agent frameworks, with full backward compatibility. Its documentation calls it a low-level orchestration framework for long-running, stateful agents. You write the graph yourself, which costs some boilerplate and gives you the most control over where state lives. Its documented tracing runs through LangSmith, LangChain's commercial product.
Google ADK became a graph runtime in its 2.0 release. ADK 1.0 shipped in May 2025 as a code-first toolkit. ADK 2.0, generally available for Python on 19 May 2026, moved from a hierarchical agent executor to a graph-based execution engine, and the Go and TypeScript versions followed in June and August. It supports Gemini natively and other models, including Claude, through connectors. It is Apache-2.0 licensed, the only 1 of the 3 that is.
Microsoft Agent Framework is the newest and has the clearest lineage. It describes itself as the direct successor to AutoGen and Semantic Kernel, built by the same teams, reached 1.0 in April 2026 with a stated commitment to long-term support, and supports Microsoft Foundry, Azure OpenAI, OpenAI, Anthropic, Ollama and others. Its workflows are graph based, with checkpoints and human-in-the-loop built in. Install agent-framework-core, not the meta-package.
Choose among them on your cloud and your team, not on features. All 3 cover the same ground. A team on Google Cloud gets the shortest path from ADK. A team on Azure gets it from Agent Framework. A team that wants to stay independent of any cloud vendor gets the most mature option in LangGraph.
These frameworks give you an agent, tools and a loop, and leave the workflow to you. That is the right shape for agents that finish in 1 request, and for teams that already run a workflow engine and only need the model-calling part.
Pydantic AI is typed from end to end. Agents declare their dependencies and their output type, and the framework validates what the model returns against that type. For a team that already uses Pydantic for validation, the learning curve is short. Durable execution comes through integrations with Temporal, DBOS, Prefect or Restate, and tracing through OpenTelemetry. It supports every major model provider.
The OpenAI Agents SDK keeps its primitives few. Its documentation describes a very small set: agents, agents as tools, handoffs and guardrails. It calls the SDK the production-ready upgrade of OpenAI's earlier experimental Swarm project. Tracing is built in. A run's state can be serialised and resumed, and fuller durable execution comes from integrations such as Temporal. It is still 0.x after 87 releases in a year.
smolagents writes its actions as code. Its CodeAgent produces Python rather than JSON tool calls, from a core of about 1,000 lines. That makes it unusually readable and unusually easy to audit. It does not document built-in durable execution, so it suits short tasks and research prototypes better than long-running production work.
Agno sits between a framework and a platform. It describes itself as a framework and runtime for building agent platforms, with agents, teams and workflows in the open-source SDK, and its durable queue is part of AgentOS, its runtime. Check which features you need come from the open-source package before committing.
These loops pair well with a separate orchestrator. If you are splitting work across several agents, the coordination patterns and their costs are covered in our multi-agent system post, and they apply whichever framework runs each agent.
The Claude Agent SDK is Claude Code as a library. Its documentation describes it as the same tools, agent loop and context management that power Claude Code, which is why it bundles the 223 MB executable. It documents Claude models only, through the Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. It has built-in permissions, human approval, sessions that resume or fork, OpenTelemetry tracing and MCP support. Its checkpointing covers file changes, not a general workflow engine.
Lock-in is a price, not a sin. A vendor SDK gives you the vendor's best loop for the vendor's model on the day it ships. What it costs is the ability to move a workload to a cheaper or better model without rewriting the agent. For a coding or file-heavy agent where Claude is the reason you are building at all, that trade is often worth making. For a customer-facing agent you may later want to move to a cheaper model, it usually is not.
Keep the vendor boundary thin either way. Put prompts, tool definitions and evaluation sets in your own code, not in framework objects. Then a framework change rewrites the loop, not the product.
The latest CrewAI could not be installed on our test machine, and pip did not say so. Asked for crewai, it installed 1.9.3, published on 30 January 2026, rather than 1.15.23, published on 28 September 2026. Asked explicitly for 1.15.23, it failed. The cause was 2 levels down: CrewAI's newer releases, including 1.15.23, require LanceDB 0.29.2 or newer, and LanceDB stopped publishing wheels for Intel Macs after version 0.25.3 in November 2025. Linux servers and Apple Silicon Macs are unaffected.
The lesson is not about CrewAI. Any framework with a deep dependency tree can resolve to a different version on a developer laptop, a CI runner and a production container, and the resolver will choose silently whenever it can find any version that installs. A team that discovers this in production has been testing a different framework from the one it ships.
4 habits prevent it. Pin the framework to an exact version. Commit a lock file generated by uv, Poetry or pip-tools. Build and run your tests in CI on the same operating system and CPU architecture as production. And print the installed framework version at startup, so a log line answers the question before anyone has to ask it.
Prefer the core package when one exists. Microsoft's core package installs 11 other packages instead of 206, and Pydantic AI's slim package 17 instead of 97. Every package you do not install is one that cannot block your upgrade.
The 2025 Stack Overflow Developer Survey asked who had used agent frameworks in the past year. Among the 3,758 developers who answered that question, Ollama led at 51.1 per cent and LangChain at 32.9, followed by LangGraph at 16.2, Vertex AI at 15.1, Amazon Bedrock Agents at 14.5, LlamaIndex at 13.3, AutoGen at 12, CrewAI at 7.5, smolagents at 3.7 and Agno at 3.4 per cent. The OpenAI Agents SDK, Pydantic AI and Google ADK were not on the list.
The broader question shows how early this still is. In the same survey, 14.1 per cent of 31,877 developers who answered said they used AI agents daily and 37.9 per cent had no plans to. LangChain's own State of Agent Engineering survey of 1,340 practitioners, run in November and December 2025, reported 57 per cent with agents in production, with quality the top barrier at 32 per cent. That survey reached LangChain's own audience, which is more likely than average to be building agents.
TypeScript teams have parallel choices. On npm on 6 October 2026, LangGraph for JavaScript was at 1.4.19, Mastra at 1.74.0, the Vercel AI SDK at 7.0.128 and the OpenAI Agents SDK for JavaScript at 0.19.0. The same 3 questions apply.
Day 1: answer the 3 questions in writing. Durable state or not, 1 model vendor or several, and how often your team can take a breaking change. The answers eliminate most of the list.
Days 2 to 4: spike 2 candidates on the same task. 1 tool, 1 human approval step and 1 deliberately failing call. Measure what our table measures on your own machines, and add what it cannot: how readable the trace is, and how long it took to make the failure visible.
Day 5: decide on the boundary, not just the framework. Whatever you pick, keep prompts, tool definitions and evaluations in your own code, so the framework is replaceable. Our AI agent development work starts from that boundary, and the single-agent design questions behind it are in our custom AI agent development guide.
The short version. On 6 October 2026, 3 graph runtimes had reached 1.0 with durable execution built in: LangGraph, Google ADK and Microsoft Agent Framework. Pydantic AI, the OpenAI Agents SDK, smolagents and Agno give you a lighter loop and leave durability to integrations. The Claude Agent SDK is the strongest single-vendor option and supports Claude only. AutoGen is in maintenance mode. Across the 10, installed size ran from 34 MB to 677 MB and stable releases from 4 to 192 a year, and the version pip gives you depends on the machine you run it on. Pin it, lock it, and choose the framework whose failures your team can read.
There is no single best framework; the right one depends on 3 questions. If agents need durable state across crashes and human pauses, choose a graph runtime that has reached 1.0: LangGraph, Google's Agent Development Kit or Microsoft Agent Framework. For a light, typed agent loop, choose Pydantic AI, the OpenAI Agents SDK or smolagents. For Claude-only agents, the Claude Agent SDK runs Claude Code's own loop.
No. Microsoft's AutoGen repository states that AutoGen is in maintenance mode, will not receive new features and is now community managed, and it directs new users to Microsoft Agent Framework. AutoGen AgentChat's last PyPI release was 0.7.5 on 30 September 2025. Microsoft Agent Framework reached 1.0 in April 2026 and describes itself as the direct successor to AutoGen and Semantic Kernel.
LangChain is a library of integrations and high-level building blocks for language model applications. LangGraph is a lower-level orchestration framework from the same company for long-running, stateful agents, modelled as a graph of steps with checkpointing and human-in-the-loop interrupts. LangGraph reached 1.0 on PyPI in October 2025. In the 2025 Stack Overflow survey, 32.9 per cent of agent framework users had used LangChain and 16.2 per cent LangGraph.
Yes, with some limits. Its documentation describes using any OpenAI-compatible endpoint, a custom model provider, or adapters for other providers, and marks 2 of those adapters as beta. The SDK is MIT licensed and was at version 0.23.1 on 6 October 2026, having published 87 releases in the preceding 12 months without reaching 1.0.
No. The Claude Agent SDK documents Claude models only, available through the Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. It packages the agent loop, tools and context management that power Claude Code, including a bundled executable of about 223 MB, and provides permissions, human approval, resumable sessions, OpenTelemetry tracing and Model Context Protocol support.
LangGraph, Google ADK 2 and Microsoft Agent Framework document built-in checkpointing or resumption of stopped runs. CrewAI documents checkpointing for crews and flows. Pydantic AI and the OpenAI Agents SDK provide durable execution through integrations with workflow engines such as Temporal and DBOS. Agno provides a durable queue through its AgentOS runtime. smolagents does not document built-in durable execution.
On 6 October 2026, on Python 3.12, installed sizes ranged from 34 MB for Microsoft Agent Framework's core package to 677 MB for its full package with all integrations. LangGraph installed 59 MB, the OpenAI Agents SDK 90 MB, Google ADK 104 MB, Pydantic AI 209 MB in full or 47 MB slim, and the Claude Agent SDK 278 MB, most of it a bundled executable.

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.
Share your details and we will talk soon.
Be the first to access expert strategies, actionable tips, and the trends actually shaping the digital world. No fluff - just practical insights delivered straight to your inbox.
Dive into our blog and stay ahead of the curve with expert perspectives, future-ready trends, and tech tips written for decision-makers and doers alike.