> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nimbleway.com/llms.txt
> Use this file to discover all available pages before exploring further.

# LlamaIndex

> Add real-time web search and cited deep research to LlamaIndex agents and workflows with the Nimble tool specs.

## Overview

The [`llama-index-tools-nimble`](https://pypi.org/project/llama-index-tools-nimble/) package brings Nimble to [LlamaIndex](https://www.llamaindex.ai/). It ships two tool specs, and an agent can hold both at once.

* **`NimbleToolSpec`**: live web search. One query, ranked results in seconds.
* **`NimbleAgentToolSpec`**: deep research on a [Web Search Agent](/nimble-sdk/web-search-agents/overview). One run, a synthesized answer with per-claim citations and confidence.
* **Citable output**: both return LlamaIndex `Document` objects with the source URLs in the document text, so the model can read and cite them.
* **Agent-ready**: `to_tool_list()` plugs either spec into any LlamaIndex agent or workflow.
* **RAG-friendly**: `Document` output drops straight into LlamaIndex indexes and query engines.

## Prerequisites

<AccordionGroup>
  <Accordion title="Python 3.10 or later">
    The package requires `>=3.10,<4.0`.
  </Accordion>

  <Accordion title="LlamaIndex core">
    `llama-index-core>=0.13.0,<0.15` installs automatically. The agent examples also need an LLM integration, such as `llama-index-llms-openai`.
  </Accordion>

  <Accordion title="nimble-python 1.2.0 or later">
    Version 0.2.0 raised the SDK floor from `>=0.16,<1.0.0` to `>=1.2.0,<2.0.0`. If other code in the same environment targets `nimble-python` 0.x, plan to update it at the same time.
  </Accordion>

  <Accordion title="NIMBLE_API_KEY">
    Required by both tool specs. Get a key from the [dashboard](https://online.nimbleway.com/settings/api-keys), free trial available. The key is passed to the SDK client rather than stored on the tool spec, and the tool's own run errors are built without it.
  </Accordion>

  <Accordion title="A Web Search Agent instance id (optional)">
    Deep research runs without one: the API provisions an agent per run. See [Agent identity](#agent-identity).
  </Accordion>
</AccordionGroup>

## Quick Start

<Steps>
  <Step title="Install">
    ```bash theme={"system"}
    pip install llama-index-tools-nimble
    ```

    Add an LLM integration to run the agent examples:

    ```bash theme={"system"}
    pip install llama-index-llms-openai
    ```
  </Step>

  <Step title="Set your API key">
    Get your key from [Nimble's dashboard](https://online.nimbleway.com/settings/api-keys), then set it as an environment variable:

    ```bash theme={"system"}
    export NIMBLE_API_KEY="your-api-key"
    ```

    Or pass it directly when constructing either tool spec:

    ```python theme={"system"}
    from llama_index.tools.nimble import NimbleToolSpec

    tool_spec = NimbleToolSpec(api_key="your-api-key")
    ```
  </Step>

  <Step title="Search the web">
    `max_results` caps how many results come back. The default is 6.

    ```python theme={"system"}
    from llama_index.tools.nimble import NimbleToolSpec

    tool_spec = NimbleToolSpec()
    documents = tool_spec.search(
        "latest developments in web data infrastructure",
        max_results=5,
    )

    for doc in documents:
        print(doc.metadata["title"], "-", doc.metadata["url"])
    ```

    ```text theme={"system"}
    Web Data Infrastructure for AI - https://nimbleway.com/...
    The State of Web Data in 2026 - https://example.com/...
    ```
  </Step>
</Steps>

## How it works

<Steps>
  <Step title="The agent receives the tool">
    `to_tool_list()` turns a tool spec into LlamaIndex tools. `NimbleToolSpec` exposes `search`, `NimbleAgentToolSpec` exposes `run`.
  </Step>

  <Step title="The model decides to call it">
    When the prompt needs live data, the model calls `search` for a quick lookup or `run` for a researched answer.
  </Step>

  <Step title="Nimble answers, the model cites">
    Results come back as `Document` objects whose text carries the source URLs, so the model writes a grounded, citable answer.
  </Step>
</Steps>

## Search

`NimbleToolSpec` wraps Nimble's [Search API](/nimble-sdk/web-tools/search) and returns each result as a `Document` carrying the page title and source URL.

### Build an AI agent

Pass the tool spec to a LlamaIndex `FunctionAgent` with `to_tool_list()`. The agent can then search the web to answer questions with current facts:

```python theme={"system"}
import asyncio
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI
from llama_index.tools.nimble import NimbleToolSpec

agent = FunctionAgent(
    tools=NimbleToolSpec().to_tool_list(),
    llm=OpenAI(model="gpt-4o-mini"),
    system_prompt="You are a research assistant. Use web search to answer with current facts.",
)

async def main():
    response = await agent.run("What has Nimble announced recently?")
    print(response)

asyncio.run(main())
```

### Citable Document output

Every result is a LlamaIndex `Document`. The title and source URL are embedded in the document text, so the agent can read and cite the source even though it never sees metadata. Each document's text has this shape:

```text theme={"system"}
<title>
URL: <source url>

<page content or snippet>
```

The title and URL are also available as metadata fields for programmatic use:

| Field | Description |
| - | - |
| `metadata["title"]` | The page title |
| `metadata["url"]` | The source URL |

### Use results in a LlamaIndex index

Because `search()` returns `Document` objects, results drop straight into a LlamaIndex index for retrieval-augmented generation:

```python theme={"system"}
from llama_index.core import VectorStoreIndex
from llama_index.tools.nimble import NimbleToolSpec

# Fetch live web results as Documents
documents = NimbleToolSpec().search(
    "web data infrastructure trends",
    max_results=10,
)

# Index them, then query with citable sources
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("Summarize the key trends.")
print(response)
```

### Search options

<AccordionGroup>
  <Accordion title="query">
    Argument to `search()`. The search query, supplied by the model when the tool is agent-driven.
  </Accordion>

  <Accordion title="max_results">
    Argument to `search()`. Maximum results to return. Type `int`, default `6`, must be at least 1. Nimble's own cap is soft, so the tool trims the list to hold the contract.
  </Accordion>
</AccordionGroup>

`NimbleToolSpec(api_key=...)` is the only constructor argument. Search depth, focus, and output format are fixed at `lite`, `general`, and markdown, and the model cannot change them.

## Deep research with Web Search Agents

`search()` is a synchronous building block: one query, one response, seconds. A [Web Search Agent](/nimble-sdk/web-search-agents/overview) run is a different capability. An autonomous agent plans, searches, reads, and cross-checks many sources, then returns a final answer with per-claim citations and confidence.

| | Search | Web Search Agent run |
| - | - | - |
| Latency | Seconds | 10 to 30 seconds at `low`, 5 to 15 minutes at the default `high`. See [Efforts](/nimble-sdk/web-search-agents/efforts) |
| Endpoint | `/v2/search` | `/v2/agents/*` |
| Output | Ranked results, one `Document` each | One `Document`: final answer plus sources, citations, confidence |
| Best for | Grounding a chat turn | Briefs, due diligence, monitoring, enrichment |

### Run one research task

`NimbleAgentToolSpec` exposes a single `run` tool. It creates the run, polls until the run reaches a terminal status, then fetches and maps the result, all inside one call.

```python theme={"system"}
from llama_index.tools.nimble import NimbleAgentToolSpec

# timeout is raised past the 300 s default, because agent runs outlast it.
agent_tool = NimbleAgentToolSpec(timeout=1800)  # reads NIMBLE_API_KEY
doc = agent_tool.run(
    "What are the leading approaches to LLM guardrails, and who builds them?",
)

print(doc.text)                             # final answer + "Sources:" list
print(doc.metadata["confidence"])           # high / medium / low / pre_existing
print(doc.metadata["claims"])               # per-claim citations
print(doc.metadata["web_search_agent_id"])  # the agent the API ran this on
```

Hand it to a LlamaIndex agent the same way as the search spec. Tell the model the tool is slow but thorough, so it reaches for it deliberately:

```python theme={"system"}
import asyncio
import os
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI
from llama_index.tools.nimble import NimbleAgentToolSpec

tool_spec = NimbleAgentToolSpec(
    agent_id=os.environ.get("NIMBLE_AGENT_ID"),
    timeout=1800,
)
agent = FunctionAgent(
    tools=tool_spec.to_tool_list(),
    llm=OpenAI(model="gpt-4o-mini"),
    system_prompt=(
        "You are a research assistant. For questions that need a researched, "
        "synthesized answer, call the Nimble research tool (it is slow but "
        "thorough) and relay its answer together with the source URLs it cites."
    ),
)

async def main():
    response = await agent.run(
        "Research: what do photography reviewers consider the best "
        "mirrorless cameras for travel, and why? Cite your sources."
    )
    print(response)

asyncio.run(main())
```

<Note>
  The model-facing tool schema is `{task, output_schema, input_data, sources}`. Agent identity, credentials, effort, and the deadline are set by the application at construction, so the model can never choose the agent or raise the cost tier.
</Note>

### Agent identity

`agent_id` is optional, and the choice decides the route:

| Configuration | Route | Behavior |
| - | - | - |
| `agent_id` omitted | `POST /v2/agents/runs` | The API provisions an agent for the run and returns its id |
| `agent_id="wsa_..."` | `POST /v2/agents/{agent_id}/runs` | The run executes on that preconfigured agent |

Either way the **returned** `web_search_agent_id`, not the configured one, is the authority for the polling and result calls that follow. It lands in the document metadata, so an auto-provisioned agent stays reachable afterwards. A returned id that differs from a configured one is rejected rather than papered over.

<Note>
  The tool is execution-only by design. It never creates, edits, or deletes agent instances, so account administration stays off the LLM surface. To pin runs to an agent, provision it in the [dashboard](https://online.nimbleway.com) or through `POST /v2/agents`, then pass its `wsa_...` id.
</Note>

### Structured output and enrichment

Three optional per-run controls shape the answer. They are passed through unchanged: the adapter never rewrites a caller's schema or source guidance.

```python theme={"system"}
doc = agent_tool.run(
    "Find the founding year and headquarters for this company.",
    input_data={"company": "Acme", "domain": "acme.example"},
    output_schema={
        "type": "object",
        "properties": {"founded": {"type": "integer"}, "hq": {"type": "string"}},
    },
    sources={"prioritize": "regulatory filings", "avoid": "press releases"},
)
```

<AccordionGroup>
  <Accordion title="output_schema">
    A JSON Schema object the answer must match. Supply it to get structured JSON back in place of prose. The document text then holds pretty-printed JSON, and `metadata["output_type"]` reads `json`.
  </Accordion>

  <Accordion title="input_data">
    An object, or a list of objects, to enrich. Each row holds known data about one entity to research. Enrichment payload only: it is never stored on the agent.
  </Accordion>

  <Accordion title="sources">
    Source guidance. `allow` and `block` take lists of source objects; `prioritize` and `avoid` take free text. Any other key is rejected, so a model filling this tool schema cannot smuggle `skill`, `use_case`, or `agent_name` through a nested dict.
  </Accordion>
</AccordionGroup>

The schema above returns this shape, prose replaced by pretty-printed JSON and the `Sources:` list kept:

```text theme={"system"}
{
  "hq": "<headquarters city>",
  "founded": <year>
}

Sources:
- <source title> — <source url> (primary)
- <source title> — <source url> (secondary)
```

### Tool spec options

Set on the `NimbleAgentToolSpec` constructor. Every parameter is optional.

<AccordionGroup>
  <Accordion title="agent_id">
    The Web Search Agent instance to run on, in the form `wsa_...`. Type `str`, default `None`, which lets the API provision an agent per run.
  </Accordion>

  <Accordion title="api_key">
    Nimble API credentials. Falls back to `NIMBLE_API_KEY`.
  </Accordion>

  <Accordion title="effort">
    `'low'`, `'medium'`, `'high'`, `'x-high'`, or the gated `'max'`. Default `None`, which omits the field so Nimble applies the agent or template default of `high`. Higher tiers research more sources and take longer. See [Efforts](/nimble-sdk/web-search-agents/efforts). This is an application setting, not a tool argument the model can raise.
  </Accordion>

  <Accordion title="gate_policy">
    Treatment for gated values: `'reject'` (default) or `'degrade'`. `max` is a coming-soon custom-budget capability. Under `reject`, `effort="max"` raises before any run is created and points at the [Nimble product team](https://www.nimbleway.com/contact). Under `degrade`, the effective tier becomes `x-high` and a warning announces the substitution, so a different tier is never billed silently.
  </Accordion>

  <Accordion title="timeout">
    Overall deadline in seconds for one `run` call, covering creation, polling, and result retrieval. Type `float`, default `300.0`. Raise it: the default is shorter than a typical run. See [Long runs and deadlines](#long-runs-and-deadlines).
  </Accordion>

  <Accordion title="poll_interval">
    Seconds between status polls, and the pause before re-attempting a transient failure. Type `float`, default `10.0`. Shorter values are intended only for tests.
  </Accordion>

  <Accordion title="agent_name, skill, use_case">
    Typed SDK hints applied to the run. `agent_name` names the run's generated agent, `skill` is a skill identifier, and `use_case` is `'research'`, `'enrichment'`, or `'dataset_building'`. All default to `None`.
  </Accordion>
</AccordionGroup>

### Results and citations

A completed run returns one `Document`. Its text is the final answer followed by a `Sources:` list, because an agent only ever sees a tool's stringified output, where document metadata is dropped. Citations have to live in the text for the model to read them.

```text theme={"system"}
# Leading Approaches to LLM Guardrails

LLM guardrail approaches are organized around a multi-layered safety pipeline
spanning the model's lifecycle [1]:

## 1. Input-Level Defenses (Prompt Screening)
Pre-processing mechanisms block or modify disallowed prompts before they reach
the LLM, typically using classifiers, keyword/regex matching, or prompt-injection
detectors [2].

...

Sources:
- Survey on LLM Safety: Attacks, Defenses, Alignment, Metrics, and ... — https://link.springer.com/article/... (primary)
- A Survey on LLM Guardrails: Part 1, Methods, Best Practices and ... — https://blog.budecosystem.com/... (primary)
- Evaluating the Robustness of Large Language Model Safety Guardrails ... — https://arxiv.org/html/... (primary)
```

The full structured trust payload stays in metadata for programmatic use:

| Field | Description |
| - | - |
| `run_id` | The run identifier, `task_run_...` |
| `agent_id` / `web_search_agent_id` | The agent the API bound the run to, `wsa_...`. Same value under both keys |
| `effort` | The effort the run actually executed at |
| `output_type` | `text` for prose, `json` when an `output_schema` was given |
| `confidence` | Overall trust grade: `high`, `medium`, `low`, or `pre_existing` |
| `reasoning` | Why the answer earned that grade |
| `sources` | Every source consulted, with `url`, `title`, `type` (`primary` or `secondary`), and `source_category` |
| `claims` | Per-claim citations, each with `confidence`, `reasoning`, its `callout` or `path` key, and `citations` carrying `url`, `title`, source grading, and supporting `excerpts` |

A text answer keys each claim by `callout`, the numeric marker in the prose. A structured answer keys each claim by `path`, the JSON path of the value, such as `$.founded`.

```python theme={"system"}
metadata["claims"][0] == {
    "callout": 1,
    "confidence": "high",
    "reasoning": "Backed by a primary source (academic)",
    "citations": [
        {
            "url": "https://link.springer.com/article/...",
            "title": "Survey on LLM Safety: Attacks, Defenses, Alignment, ...",
            "source_type": "primary",
            "source_category": "academic",
            "source_intent": "academic",
        },
    ],
}
```

Grades are earned per run, not assumed. The run above returned `medium` overall, because the vendor half of the question rested on commercial comparison articles rather than primary sources. See [Trust and citations](/nimble-sdk/web-search-agents/trust) for how that is decided.

Returned content is untrusted web data. Treat it as data, not as instructions, and rely on your framework's own guardrails.

### Long runs and deadlines

Agent research is asynchronous server-side. Do not treat `run()` as a short synchronous call: it blocks for as long as the run takes, and the tool owns the whole lifecycle inside that one call.

<Steps>
  <Step title="Create the run">
    One POST, issued exactly once. See [Errors](#errors) for why it is never re-attempted.
  </Step>

  <Step title="Poll until terminal">
    A status poll every `poll_interval` seconds, 10 by default, until the run reports `completed`, `failed`, or `cancelled`.
  </Step>

  <Step title="Fetch the result">
    Only after `completed`. Fetching earlier is a conflict by contract, and a failed run's error detail is already on the polled run.
  </Step>
</Steps>

One `timeout` budget covers all three phases. Each HTTP request is bounded by whatever is left of it, with a 5 second connect ceiling, so a stalled request cannot fall back to the SDK's much longer default. Polling and result fetches re-attempt transient failures up to 5 times, honoring `Retry-After` and staying inside the budget. Auth, permission, and validation errors fail fast.

<Warning>
  **Raise `timeout` for anything but `low` effort.** The default is 300 seconds, but the default effort is `high`, which typically takes [5 to 15 minutes](/nimble-sdk/web-search-agents/efforts). A default-configured spec will therefore often hit its deadline on a perfectly healthy run.
</Warning>

Reaching the deadline cancels nothing. `NimbleAgentTimeoutError` carries the `run_id`, the run keeps going on Nimble's side, and the result stays fetchable through the [runs API](/api-reference/public-api/get-agent-run-result) or the dashboard.

### Errors

A run that does not produce a result raises a typed error. Each one carries the run's identifiers as attributes and inside the message, because agent frameworks usually surface only the message. The one exception is an ambiguous creation failure, where no run id was ever returned.

| Error | Meaning |
| - | - |
| `NimbleAgentTimeoutError` | Not terminal within `timeout`. The run may still complete server-side |
| `NimbleAgentRunFailedError` | The run terminated as `failed`. Carries the server's error message |
| `NimbleAgentRunCancelledError` | The run terminated as `cancelled` |
| `NimbleAgentProtocolError` | Unknown status, malformed result, response identity mismatch, or persistent polling and result errors. The SDK exception is chained |
| `NimbleAgentCreateAmbiguousError` | Run creation failed without a definite outcome. `run_id` is `None`, and `status_code` carries the HTTP status when there was one |

All five subclass `NimbleAgentRunError`, so one `except` clause catches every run failure:

```python theme={"system"}
from llama_index.tools.nimble import (
    NimbleAgentRunError,
    NimbleAgentTimeoutError,
    NimbleAgentToolSpec,
)

# x-high typically takes 15 to 30 minutes, so the deadline allows for that.
agent_tool = NimbleAgentToolSpec(effort="x-high", timeout=2400)

try:
    doc = agent_tool.run("Map the EU AI Act enforcement timeline.")
except NimbleAgentTimeoutError as error:
    # Not a failure. Persist the ids however your app stores state, then
    # fetch the run later. `save_run_for_later` here is your own function.
    save_run_for_later(error.run_id, error.agent_id)
except NimbleAgentRunError as error:
    print(error.status, error.run_id, error.agent_id)
    raise
```

<Note>
  `NimbleAgentCreateAmbiguousError` is the one error not to retry. Creation is sent once and never re-sent, so a run may already be underway with no id returned. A function-calling model reads a bare timeout as an ordinary transient, so exclude this error from any automatic tool-retry policy. Look for the earlier run in the [dashboard](https://online.nimbleway.com) instead. Requests that were definitely rejected, such as a bad key or a validation error, are safe to call again once fixed.
</Note>

## Limitations

* **Search and Web Search Agent runs** ship in this package. Extract, Map, and Crawl are not exposed as tool specs. Use the [Nimble Python SDK](/nimble-sdk/sdks/python).
* **One blocking call per run.** There is no start-now, collect-later split, and no run event streaming (SSE). A run that outlives `timeout` is recovered through the runs API, not through this tool.

## Related features

Easily confused with this package:

| Feature | Where it lives |
| - | - |
| Web Search Agents, `/v2/agents/*` | This page. Current deep-research agents |
| Extract Templates, `/v2/extract/templates` | The [Extract Template guide](/nimble-sdk/web-tools/extract/template). The old `/v1/agent` site scrapers were renamed to these, and are unrelated to Web Search Agents |
| LlamaHub listing | The package carries LlamaHub registry metadata for both tool specs. No effect on how you install or import |
| LlamaParse | LlamaIndex's own document parser. Not a Nimble product and not related to these tool specs |

## Resources

<CardGroup cols={2}>
  <Card title="PyPI Package" icon="python" href="https://pypi.org/project/llama-index-tools-nimble/">
    `llama-index-tools-nimble` on PyPI.
  </Card>

  <Card title="GitHub Repository" icon="github" href="https://github.com/Nimbleway/llama-index-tools-nimble">
    Source, README, and runnable examples.
  </Card>

  <Card title="v0.2.0 Release" icon="tag" href="https://github.com/Nimbleway/llama-index-tools-nimble/releases/tag/v0.2.0">
    Deep research agents and agentless runs.
  </Card>

  <Card title="Web Search Agent" icon="robot" href="/nimble-sdk/web-search-agents/overview">
    The deep-research capability behind the `run` tool.
  </Card>

  <Card title="Search API" icon="magnifying-glass" href="/nimble-sdk/web-tools/search">
    Nimble's underlying search capability.
  </Card>

  <Card title="Nimble Python SDK" icon="python" href="/nimble-sdk/sdks/python">
    Direct access to Extract, Map, Crawl, and agent administration.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.