Profile, evaluate, and tune agents built with LangChain, CrewAI, and other frameworks.
GitHub Stars
2.7K
Contributors
124
npm / Week
—
PyPI / Month
88.5K
7%The NVIDIA NeMo Agent Toolkit is an open-source Python library designed to connect, profile, evaluate, and tune AI agents across heterogeneous environments. Maintained by NVIDIA under the Apache 2.0 license, it operates not as a ground-up replacement for existing libraries, but as an umbrella layer that sits beside tools like LangChain, CrewAI, LlamaIndex, Google ADK, and Microsoft Semantic Kernel. Where most frameworks focus on prompt chaining or graph topology, NeMo Agent Toolkit addresses the operational gap that appears when multi-agent systems transition from local prototypes to production infrastructure.
The framework occupies a distinct intersection between observability, agent runtime orchestration, and automated workflow optimization. Rather than forcing teams to rewrite existing application logic, it wraps existing agents in a unified interface, instrumenting execution paths and token consumption out of the box. With more than 2,600 GitHub stars, 120 contributors, and nearly 100,000 monthly PyPI downloads, it has emerged as a key utility for enterprise teams trying to bring rigorous profiling and benchmarking to otherwise black-box agent loops.
Behind its design philosophy is NVIDIA's emphasis on hardware efficiency and inference optimization. As multi-agent workflows chain dozens of LLM calls together, latency, inference costs, and tool failure rates compound rapidly. NeMo Agent Toolkit provides the instrumentation necessary to isolate bottlenecks down to individual tokens, then pairs that profiling with automated prompt and hyperparameter optimizers to maximize task throughput on modern compute stacks.
What the framework gives you out of the box, in plain language.
The jobs this framework is best suited for.
Profile an existing LangChain or CrewAI agent to see where time and tokens go before it reaches production.
Run evaluation datasets against each new version of an agent and compare the scores before shipping.
Teams running NIM models or Dynamo that want agents tuned for speed and lower cost at scale.

Side-by-Side
Add a second or third framework and see stars, downloads, and capabilities lined up next to each other.
An agent framework is the code your team uses to wire large language models into tools, memory, and human checkpoints. It is the connective tissue between an LLM call and a real task, like answering a support ticket or running a multi-step research workflow.
NVIDIA NeMo Agent Toolkit ships under the apache-2 license. The source code lives on GitHub, so you can read it, fork it, and run it on your own infrastructure if your team prefers self-hosting.
NVIDIA NeMo Agent Toolkit is primarily a Python project. Pick a framework that matches the language your team already ships in. The cost of a stack switch is almost always higher than the difference between two frameworks.
The NeMo Agent Toolkit architecture centers on a plugin-based runtime driven by function composability and declarative YAML definitions. At its core, every tool, agent, and memory store is modeled as a callable component registered via entry points and Python decorators. This lets developers treat custom agents built in CrewAI or LangChain as modular nodes inside a larger NVIDIA NeMo Agent Toolkit workflow.
Control flow can be managed imperatively in Python or declaratively through YAML files. In the declarative model, engineers define LLM clients, tool bindings, orchestration logic, and middleware in a single structured configuration. Under the hood, the toolkit uses a middleware architecture (FunctionMiddleware, HITLMiddleware) that intercepts tool calls and agent transfers. This makes it possible to inject capabilities like human-in-the-loop approvals, structured logging, token accounting, and streaming without modifying agent business logic.
The runtime exposes multiple interaction planes. A configured workflow can be tested locally using the nat CLI, deployed as a standard REST service, or served directly using modern agentic communication protocols. Specifically, the toolkit supports the Model Context Protocol (MCP) as well as Agent-to-Agent (A2A) specifications, acting as either a client or a server. This design lets an enterprise deploy one NeMo Agent Toolkit instance that orchestrates external MCP tools while exposing its own specialized agents to other systems across the network.
NeMo Agent Toolkit packages deep profiling, continuous evaluation, and cross-framework runtime controls into a unified toolset.
Teams adopt NeMo Agent Toolkit when simple agent loops outgrow basic logging and start consuming unpredictable computing resources.
In complex setups where multiple agents pass tasks back and forth, pinpointing why a run takes 30 seconds can be difficult. Engineering teams use NeMo Agent Toolkit to profile existing LangChain or CrewAI pipelines before release. The profiler reveals whether latency stems from an inefficient retrieval step, an overly verbose tool schema, or redundant reasoning loops, allowing developers to target the exact bottleneck.
Before pushing updates to production agents, teams integrate NeMo Agent Toolkit into their CI/CD pipelines. By running deterministic evaluation suites against staging branches, teams catch reasoning regressions, tool hallucination spikes, or prompt drifts before code lands in production.
Enterprise deployments operating on private infrastructure often serve models through NVIDIA NIM microservices. NeMo Agent Toolkit directly integrates with NIM, allowing teams running models like Llama 3 or Nemotron to tune system prompts, context lengths, and concurrency settings specifically for their hardware configuration, lowering total cost of ownership.
The library is distributed via PyPI under the package name nvidia-nat. You can install it using standard Python package managers:
1pip install nvidia-nat2# or using uv3uv pip install nvidia-nat
Building a minimal workflow typically involves defining an agent configuration in YAML, which binds models and tools together, and then executing it via Python or the nat command-line interface.
1# workflow.yaml2models:3 primary:4 provider: nim # or openai5 model: meta/llama-3.3-70b-instruct6 api_key: ${NVIDIA_API_KEY}78tools:9 - name: search_docs10 module: my_tools.search1112agents:13 - name: support_agent14 model: primary15 tools: [search_docs]16 instructions: "Answer user queries accurately using available search tools."
You can then run or evaluate the defined workflow directly:
1from nat.runtime import WorkflowRunner23# Load and execute the declarative configuration4runner = WorkflowRunner.from_config("workflow.yaml")5response = runner.run(input_text="Summarize our enterprise SLA policies.")67print(response.output)8# Access detailed profiling data generated automatically during the run9print(f"Tokens consumed: {response.profile.total_tokens}")10print(f"Execution time: {response.profile.latency_seconds}s")
To run this pipeline, you need an API key for your chosen LLM endpoint (such as an NVIDIA API key for NIM or an OpenAI key), a Python 3.10+ environment, and any third-party observability keys (such as LangSmith or Phoenix) if you plan to export traces. Comprehensive guides, API references, and sample workflows are maintained in the official NVIDIA NeMo Agent Toolkit documentation.
Evaluating the NeMo Agent Toolkit requires understanding where its boundaries lie relative to standard orchestrators.
LangChain and CrewAI are orchestration frameworks designed for structuring prompts, defining memory buffers, and managing multi-agent role-playing loops. NeMo Agent Toolkit is not a competing orchestration model; it is an instrumentation, evaluation, and operational layer that wraps around them.
If your goal is simply to build a multi-agent chat loop with high-level Python abstractions, CrewAI or LangChain will get you moving faster. However, when you need token-level profiling across disparate frameworks, automated hyperparameter and prompt optimization, or unified MCP endpoint exposure, NeMo Agent Toolkit provides the necessary tooling without requiring you to discard your existing agent code.
Observability platforms like Langfuse and Arize Phoenix focus on collecting traces, tracking production logs, and providing web dashboards for debugging. NeMo Agent Toolkit includes native tracing and even exports to these platforms, but it extends further into runtime execution and active tuning.
Beyond capturing telemetry, NeMo Agent Toolkit includes an optimization engine that modifies prompts, searches configuration spaces, and acts as a runtime router supporting MCP and A2A. If you only need a UI to inspect traces, a dedicated observability platform suffices. If you need to actively optimize performance, orchestrate agents across different libraries, and unify tool execution under a single runtime, NeMo Agent Toolkit is the more complete technical layer.
Trace whole workflows down to single tokens. Export traces to LangSmith and other OpenTelemetry tools.
Run offline evaluations on your workflows, then search for the best prompts and model settings automatically.
Describe tools, models, and agents in one config file. Run it with the nat CLI or serve it as a REST, MCP, or A2A endpoint.