There is a point in an agent project where the framework starts costing more than it returns. The symptom is predictable: you are reading framework source to find out why a tool call did not fire, or pinning versions because a minor release changed a node signature, or explaining a runnable stack trace to someone who just wanted to know which prompt ran.
Build directly on raw model calls and structured output when you have one to three tools and simple state; reach for LangGraph or CrewAI only when you need durable checkpointing, human-in-the-loop interrupts, or multi-step routing that a raw tool loop cannot maintain.
This post works through the tradeoffs in order: the raw tool loop, the overhead frameworks add, native tool calling across OpenAI, Anthropic, and Gemini, structured output with Pydantic and related libraries, Anthropic’s workflow-versus-agent distinction, the cases where frameworks genuinely win, and a decision table. For a wider survey of orchestration patterns, see NiteAgent’s guide to agent architecture. The reasoning-then-acting loop traces to the ReAct paper, and tool-exposure standards such as the Model Context Protocol are orthogonal to whether you run a graph runtime.
The Simplest Solution Wins: When a Raw Tool Loop Is Enough
The decision to skip a framework starts with the number of tools and state, because Anthropic’s guidance is explicit: find the simplest solution possible and increase complexity only when needed, which maps directly to a one-to-three-tool raw loop (Anthropic Building Effective Agents).
The raw loop is small enough to hold in your head. You pass tool schemas to the model, the model returns a call, you execute it, you append the result, and you repeat until the model returns a message with no tool calls. That is the entire control flow. No node registry. No edge conditions. No state-channel abstraction.
Fifty lines is a realistic size for a single-tool agent, including argument parsing, error handling, and an iteration cap. When your tool count is one to three and your state is a message array plus a few session fields, a framework is indirection. Every debugging session ends in framework internals rather than your code, and every upgrade risks behavior changes in a layer you did not write.
The loop is also testable as a plain function. Mock the model client, return a canned tool call, assert that your executor received the right arguments and that the result landed back in the message list. That property matters more than configurability when the goal is to ship a working agent.
A sequential workflow — classify, then extract, then act — is likewise just three calls with typed outputs between them. If you can describe the whole system on a whiteboard without referring to a runtime, you do not need one.
The Hidden Overhead Behind Agent Frameworks
Framework abstractions create real costs in dependency bloat, debugging difficulty, and pre-1.0 API churn; the primary production evidence is Octomind’s public write-up on removing LangChain in production and moving to raw model tool calls, discussed at length by practitioners (Hacker News discussion).
Three costs recur. First, dependency bloat: pulling in an agent framework often brings a large transitive tree, so your lockfile grows faster than your agent logic. Second, debugging difficulty: an exception surfaces inside an abstraction layer, and the stack trace tells you about nodes and runnables before it tells you which prompt hit which model. Third, pre-1.0 churn: the framework’s own release history is the most direct evidence available, since breaking changes land across minor versions (LangChain releases).
Surface area compounds the churn. The LangChain repository spans many packages and integrations, and the python.langchain.com docs describe a set of abstractions considerably wider than most agents use. Every abstraction is a thing your team must learn, version, and reason about during an incident.
The Octomind case is the concrete anchor. The maintainers describe running LangChain in production for over 12 months, from early 2023 into 2024, before removing it and moving to direct model calls with their own tool loop, as recounted in the thread (Hacker News discussion). Read the thread in full before copying either side of the argument; the counterarguments about shared conventions and onboarding are part of the record.
The Raw Tool-Call Loop: OpenAI, Anthropic, and Gemini
All three major providers expose the same native pattern—send tool schemas, receive a structured call, execute it, return the result—and the canonical OpenAI function-calling guide documents this loop directly as model-generated JSON arguments that your code executes and appends (OpenAI function calling).
import json
from openai import OpenAI
client = OpenAI()
TOOLS = [{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up an order by ID.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
"additionalProperties": False,
},
"strict": True,
},
}]
def execute(name, args):
return {"order_id": args["order_id"], "status": "shipped"}
messages = [{"role": "user", "content": "Where is order A-1042?"}]
for _ in range(8): # hard iteration cap
response = client.chat.completions.create(
model="your-model-id", messages=messages, tools=TOOLS
)
msg = response.choices[0].message
messages.append(msg)
if not msg.tool_calls:
break
for call in msg.tool_calls:
result = execute(call.function.name, json.loads(call.function.arguments))
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
The shape is provider-agnostic. Swap the client and the message adapter and the control flow is identical (OpenAI Python SDK, Anthropic Python SDK).
Two details matter in production. Cap iterations, because a malformed call that raises inside your executor can loop forever if the error is fed back without a limit. And keep parallel tool calls: OpenAI and Anthropic both allow multiple calls in one turn, and executing them concurrently is an asyncio.gather away.
Anthropic documents the same request/response cycle with tool_result blocks (Anthropic tool use, Claude tool use overview), and Google documents function declarations plus the response parts you execute (Gemini function calling). None of them require a graph.
Structured Output Without a Framework: Pydantic, Instructor, Outlines, and Provider Schemas
Structured output without a framework remains fully available through typed schemas and constrained decoding, because Pydantic turns typed models into JSON Schemas for validation and model-dump, while OpenAI’s strict json_schema mode and Instructor’s validation retries add reliability without adopting an agent graph (Pydantic).
OpenAI’s strict mode constrains generation to a schema you supply (OpenAI structured outputs), Gemini offers an equivalent response schema (Gemini structured output), and Instructor wraps your existing client with a response_model argument plus a retry loop on validation failure (Instructor, Instructor on GitHub).
Self-hosted models have constrained decoding too. Outlines compiles schemas into generation constraints (Outlines), and llama.cpp accepts GBNF grammars (llama.cpp grammars). The pattern is identical everywhere: define the shape, push the decoder toward it, validate the result. Pydantic handles validation, serialization, and schema generation as a standalone library (Pydantic library). For a deeper treatment, see NiteAgent’s structured outputs guide.
from pydantic import BaseModel, Field
import instructor
from openai import OpenAI
class Triage(BaseModel):
category: str = Field(description="billing | bug | other")
priority: int = Field(ge=1, le=5)
summary: str
client = instructor.from_openai(OpenAI())
result = client.chat.completions.create(
model="your-model-id",
response_model=Triage,
max_retries=3,
messages=[{"role": "user", "content": "Refund not received for invoice 88."}],
)
print(result.category, result.priority)
Validation retries are the piece teams skip. A value can be schema-valid and semantically wrong, and re-prompting with the validation error catches most of that without any orchestration layer.
Workflows Are Not Agents: Anthropic’s Definitions
Anthropic’s primary guidance separates workflows, where LLMs and tools follow predefined code paths, from agents, where LLMs dynamically direct their own tool use, and it explicitly recommends finding the simplest solution possible before adding complexity (Anthropic Building Effective Agents). This section applies that distinction to the framework decision.
The named workflow patterns are prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer. Each is expressible as plain code. Prompt chaining is sequential calls with gates between them. Routing is a classifier call that picks a branch. Parallelization is concurrent calls plus a join. Orchestrator-workers is a planner call followed by fan-out. Evaluator-optimizer is a generate call plus a critique call in a bounded loop.
Run the classification test on your own system: if you can draw the control flow as a flowchart before the model runs, you have a workflow. Workflows do not need a node registry, an edge-condition DSL, or a graph runtime. A large share of systems described as agents in production are workflows with a tool loop bolted on, and they are usually better off described that way.
Where Frameworks Earn Their Place: Persistence, Human-in-the-Loop, and State
Frameworks earn their place only when the raw loop must become durable, resumable, or auditable, because LangGraph documents persistence checkpoints for resume and replay, built-in interrupts for human approval, and shared graph state across typed nodes (LangGraph persistence). These are non-trivial features to hand-build.
Persistence is the first real requirement. A job that must survive a restart, resume mid-graph, or replay a failed run needs checkpoints written after each step, which the persistence docs describe directly, along with durable execution semantics for runs that outlive a single request (LangGraph durable execution). Doing that by hand means designing your own checkpoint table, serialization format, and idempotency rules for every tool.
Human-in-the-loop is the second. The interrupt primitive pauses a graph, surfaces the pending action, and resumes from a checkpoint once a human responds, including inspecting and editing the pending tool call (LangGraph interrupts). You can build this with a status column and a queue, but you own the resume semantics.
Multi-step routing with shared state is the third. When many tools and branches read and write the same typed state, the graph model fits better than a message array passed around and mutated (LangGraph graph API, LangGraph overview, LangGraph repository).
CrewAI occupies a different point: role-based crews where agents are defined by roles, goals, and tasks and then coordinated (CrewAI, CrewAI docs, CrewAI crews). That pays off when a team wants a declarative definition of who does what.
Decision Table: Raw Calls vs LangGraph/CrewAI
A decision table removes guesswork by aligning the implementation path with criteria such as tool count, durability, human approval, team size, debug style, and structured-schema ownership, because each row reflects the sourced tradeoffs already established in the previous sections (Anthropic Building Effective Agents).
| Criterion | Raw model calls + structured output | Framework (LangGraph/CrewAI) |
|---|---|---|
| Number of tools | 1–3 tools; single-level tool loop | Many tools; multi-step / nested / parallel routing |
| State / durability | Stateless or trivial persistence in your own DB | Durable checkpointing, resume, time travel, long-running jobs |
| Human-in-the-loop | Simple pause/approve in code | Built-in interrupts, resume from checkpoint, audit/edit tool calls |
| Team size | Solo/small team; minimal abstraction ownership | Larger team wanting shared graph/agent abstractions |
| Debugging | Direct code path, SDK logs, raw tool messages | Requires framework tracing/visualization to unwrap abstraction |
| Schema / structured output | Provider structured output, Pydantic, or Instructor | Frameworks may wrap or constrain schema; extra layer |
Read the table left to right, not as a scoreboard. The rows are independent: you can need durable checkpointing with two tools, and you can have twelve tools with no persistence requirement at all. The rows that most often flip a decision are durability and human-in-the-loop, because both are expensive to retrofit onto a loop that already shipped.
Schema ownership is the row most often misused. Provider structured output and Pydantic already cover that layer, so a framework wrapping a schema adds a layer rather than removing one; the constraint becomes part of the abstraction you are buying. For a side-by-side of the two frameworks against the OpenAI SDK, see NiteAgent’s LangGraph vs CrewAI comparison.
Desk-research disclosure: This article was built from documented provider guides, framework docs, engineering blog posts, and public forum threads only; no hands-on testing, private benchmarks, or original performance measurements were performed for this piece.
FAQ
The FAQ section answers the five production questions practitioners ask before skipping a framework, and it uses only the primary-source claims already cited above—not new benchmarks or invented statistics—so each answer remains checkable against the original documentation (Anthropic Building Effective Agents).
Is LangGraph always overkill for agents?
No. It is overkill when your agent has one to three tools, no durable state, and no human approval step, because a raw tool loop covers that case with less code. It earns its place when runs must resume after a crash, pause for human review, or traverse many typed nodes sharing state. The answer depends on requirements, not on framework popularity.
Does skipping CrewAI or LangGraph mean giving up structured outputs?
No. Structured output lives at the provider and library layer, not the orchestration layer. OpenAI’s strict json_schema mode, Gemini’s response schema, Pydantic models, and Instructor’s response_model with validation retries all work with a plain client. Frameworks can add schema conveniences, but they are not the source of schema enforcement. The tradeoff is that you own the retry policy and validation error handling yourself.
Which provider should I choose for strict raw structured output?
Choose based on your existing stack and the decoding guarantees you need. OpenAI documents strict json_schema adherence; Gemini documents response schemas through the same API surface; Anthropic’s tool use returns typed input objects you can validate with Pydantic. For self-hosted models, Outlines and llama.cpp grammars apply constraints at the decoder. The orchestration advice does not change with the provider, so test schema adherence on your own prompts.
When should I actually adopt LangGraph or CrewAI?
Adopt when at least one hard requirement appears: durable checkpointing and resume, human-in-the-loop interrupts with the ability to inspect and edit pending tool calls, or many tools routed across shared typed state. Adopt CrewAI when a declarative role-and-task crew model matches how your team plans work. If none of those apply, the framework is optional and the raw loop is cheaper to own.
Can Pydantic replace a framework for schema handling?
Pydantic covers validation, serialization, and JSON Schema generation, which is most of the schema work an agent needs. Paired with a client and a retry loop, it handles structured output end to end. It does not provide checkpointing, interrupts, or routing, so it replaces the framework’s schema layer only, not its orchestration layer. Those are separate concerns and should be evaluated separately.
The Bottom Line
The bottom line is a verdict, not a repeat summary: raw model calls plus structured output win for one-to-three-tool stateless agents, and adoption of LangGraph/CrewAI is justified only by durable checkpointing, human-in-the-loop interrupts, or multi-step routing that raw loops cannot maintain (Anthropic Building Effective Agents).
Raw calls plus structured output are the correct default for small, stateless agents. The loop is short, the schema layer is external, and every failure traces back to code you wrote. Frameworks do not lose on capability in that range; they lose on cost, surface area, and debugging distance.
Adopt LangGraph when a run must survive a restart, pause for a human, or move through many typed nodes sharing state. Adopt CrewAI when a role-based crew definition matches how your team plans work. Those are specific, testable requirements. If you cannot name one of them for your system, you are paying framework costs for nothing.
Migration is not a one-way door. The Octomind account of removing LangChain after more than a year in production shows the raw loop is a viable destination (Hacker News discussion), and the reverse exists too: a well-factored tool executor can be wrapped as a node later, once a durability requirement actually arrives.
How This Guide Was Built
This E-E-A-T block documents the desk-research method behind the guide: primary provider docs, framework docs, engineering blog posts, and public forum threads were cited inline, and no hands-on performance testing or benchmark data was collected for this article (OpenAI function calling).
Three source classes were used. Provider documentation: OpenAI’s function calling and structured outputs guides, Anthropic’s tool use pages, and Google’s Gemini function calling and structured output docs. Framework documentation: LangGraph’s persistence, interrupts, durable execution, graph API, and overview pages, plus CrewAI’s docs and crew concepts. Engineering and community material: Anthropic’s Building Effective Agents post, the ReAct paper, and public forum threads including the Hacker News discussion of removing LangChain in production.
📖 Related Reads
- CodeIntel Log — code quality, debugging, and software engineering benchmarks
Cross-links automatically generated from NiteAgent.
← Back to all posts


