Skip to content
All library documents

Improving Stock Selection with Large and Small Order Flow Factors

Article Lumibot

Summary

This Chinese-language research summary studies stock-selection signals from large and small investor order flows. It reports that the two flows are negatively related and that normalized net flows have opposite associations with subsequent returns: large-order flow is positive, while small-order flow is negative. The authors attribute these patterns to large traders’ informational advantage and a crowding-out effect on smaller traders.

The summary compares normalization by turnover, total buys plus sells, and absolute net flow, reporting stronger information ratios for the absolute-net-flow version. It then uses cross-sectional regression to remove contemporaneous return effects from flow intensity; the resulting residual factors show improved reported information coefficients and information ratios. Similar residualization is applied to reversal, with reported long-short information ratios above the conventional reversal factor. The supplied text is an abstract rather than the underlying paper, so it omits sample construction, test period, transaction costs, and detailed robustness evidence; the reported results should not be treated as proof of live tradability.

Key ideas

  • Large-order and small-order flows are reported to have negatively related cross-sectional signals.
  • Normalizing net flow by absolute net flow is reported to outperform other tested scaling choices.
  • Cross-sectional regression is used to remove return effects from flow intensity.
  • Residual flow and reversal factors show stronger reported information ratios than their baseline versions.
  • The abstract does not provide enough detail to assess costs, sample design, or live performance.

Tags

Full text
# agents


Build AI Trading Agents in Python with LumiBot
==============================================

.. meta::
   :description: Build AI trading agents in Python with LumiBot. Start with a complete agentic backtest, then explore stock teams, macro research, and options agents.

Build AI trading agents in Python inside a LumiBot strategy. Agents can inspect
market evidence, use research tools, and submit orders through the strategy's
broker. You choose when they run and which agents can trade.

Choose your first workflow
--------------------------

* **Build your first agent:** :doc:`agents_quickstart` has installation, model credentials, daily data, and a complete researcher-and-trader backtest.
* **Trade stocks:** start with :doc:`a large-cap stock team <agents_example_bull_vs_bear_ai_stock_trading_bot>` or :doc:`opening range breakout <agents_example_opening_range_breakout_ai_trading_bot>`.
* **Explore macro teams:** inspect :doc:`the idea-meritocracy example <agents_example_ray_dalio_idea_meritocracy>` and :doc:`FRED/ALFRED data setup <macro_data>`.
* **Trade options:** :doc:`agents_example_iron_condor_ai_trading_bot` explains option-chain evidence, four-leg orders, and data limitations.

Compare prerequisites and evidence in :doc:`agents_examples` before choosing a
strategy. Start with regular stocks or ETFs; leveraged instruments and short-dated
options are advanced examples.

Run a hosted example
--------------------

The sector-pod and macro-team pages link to their regular and leveraged BotSpot
marketplace variants. BotSpot provides the hosted backtest, broker-connection,
artifact, and scheduling workflow around LumiBot. See :doc:`botspot_mcp` for
access from an AI coding assistant. Model, data, broker, and BotSpot plan
requirements depend on the example.

**Building a product on LumiBot?** :doc:`PARTNERSHIPS` explains funded
integrations, maintenance, developer tutorials, and strategic collaboration.

.. toctree::
   :maxdepth: 1

   agents_flows
   agents_builtin_tools
   agents_browser_tools
   agents_canonical_demos
   agents_observability
   agents_memory
   agents_notifications

Runtime concepts
----------------

Create agents in ``initialize()`` and call them from strategy lifecycle methods
such as ``on_trading_iteration()``. Separate research-only agents from those
allowed to submit orders. The :doc:`quick start <agents_quickstart>` is the
complete first-run example; the snippets below explain individual capabilities.

* Built-in tools provide market/account evidence and order workflows.
* ``@agent_tool`` exposes a Python function and its contract to the agent.
* Compatible MCP servers supply external tools; their authentication, schemas,
  and historical-data behavior must be checked for the intended task.
* Replay caching can reuse eligible prior agent results. A cache hit is not a
  new model decision or independent validation of a strategy.
* A broker-backed runner and a backtest runner can use the same strategy class,
  but still require different data, credentials, and execution configuration.

Recommended team architecture
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

For fully agentic trading, we recommend **two or more agents**: one or more
research agents and a **dedicated trading and risk agent** that alone can
submit or change orders. Ten researchers and one trader are just as valid as
one researcher and one trader. This is a recommendation, not a framework requirement;
LumiBot does not impose a fixed team size.

When risk rules must be mechanically fixed, keep execution and limits in
**deterministic Python** instead. A hybrid can also use agents for research and
Python for execution. Choose the ownership model deliberately, test it, and do
not give research-only agents trading permission.

Verification and historical limits
----------------------------------

Market tools use the strategy clock where supported. An LLM may nevertheless
know facts from after a historical window. Inspect source timestamps, revisions,
orders, and artifacts; a successful run does not establish profitable trading.

Release verification exercises the actual Strategy, AgentManager, built-in
tools and backtesting broker. Market observations and research responses are
fixtures; actor and judge calls use real models. Execution scenarios require a
broker-observed simulated fill, not merely an order claim in model prose.
Historical research fixtures also exercise MCP schema discovery and as-of binding.

How ``@agent_tool`` Works
-------------------------

The ``@agent_tool`` decorator is the primary way to give your AI agent access to external data. It wraps a Python method as a callable tool that the agent can invoke during its reasoning loop.

**Key feature: automatic source code inclusion.** When you decorate a method with ``@agent_tool``, LumiBot automatically includes the function's source code in the tool description sent to the AI. This means the AI can see all parameters, default values, and implementation details without you having to describe them manually. Write a clear docstring with an ``Args`` section, and the AI will understand how to call your tool correctly.

The introductory macro examples on this page use Lumibot's built-in FRED tools. Those tools require ``FRED_API_KEY`` and use official FRED/ALFRED realtime parameters so backtests do not accidentally see future macro revisions.

.. code-block:: python

    @agent_tool(
        name="search_news",
        description="Search recent stock market news from Alpaca.",
    )
    def search_news(
        self, start: str = "", end: str = "", symbols: str = "", limit: int = 10
    ) -> dict:
        """Call the Alpaca News API for historical news.

        Args:
            start: Start timestamp in ISO format
            end: End timestamp in ISO format
            symbols: Comma-separated stock symbols to filter by
            limit: Maximum number of articles to return
        """
        # The AI sees this entire function body automatically
        resp = requests.get("https://data.alpaca.markets/v1beta1/news", ...)
        return resp.json()

When you pass custom tools via ``tools=[self.my_tool]``, they are added **alongside** the default built-in tools. You only need to list your custom tools -- built-in tools are always included.

The one exception is outbound network access. ``http_request``, ``rss_fetch``, and the ``browser_*`` tools are off by default; pass ``allow_network=True`` to the agent that fetches pages. See :ref:`the network permissions section <agents-network-permissions>`.

External Data Patterns
----------------------

**Pattern 1: @agent_tool wrapping a REST API (recommended)**

This is the primary and recommended approach. It works reliably in both backtests and live trading because you control the HTTP call directly.

.. code-block:: python

    import os
    import requests
    from lumibot.components.agents import agent_tool

    @agent_tool(
        name="get_stock_bars",
        description="Get historical daily price bars for a stock from Alpaca.",
    )
    def get_stock_bars(
        self, symbol: str, start: str = "", end: str = "", limit: int = 30
    ) -> dict:
        """Get historical OHLCV bars from the Alpaca market data API.

        Args:
            symbol: Stock ticker symbol (e.g., TQQQ, SPY, QQQ)
            start: Start date in YYYY-MM-DD or ISO format
            end: End date in YYYY-MM-DD or ISO format
            limit: Maximum number of bars to return
        """
        api_key = os.environ.get("ALPACA_API_KEY", "")
        api_secret = os.environ.get("ALPACA_API_SECRET", "")
        headers = {"APCA-API-KEY-ID": api_key, "APCA-API-SECRET-KEY": api_secret}
        params = {"timeframe": "1Day", "limit": limit, "sort": "desc"}
        if start:
            params["start"] = start
        if end:
            params["end"] = end
        resp = requests.get(
            f"https://data.alpaca.markets/v2/stocks/{symbol}/bars",
            headers=headers, params=params, timeout=15,
        )
        return resp.json()

This pattern works with any REST API -- Alpaca, FRED, Alpha Vantage, or your own internal services. All four demo strategies use this approach.

**Pattern 2: MCP server via URL (for live trading or compatible servers)**

If you have a compatible MCP server, you can connect it by URL. This is useful for live trading scenarios or when a third-party provides a dedicated MCP server.

.. code-block:: python

    from lumibot.components.agents import MCPServer

    MCPServer(
        name="my-data-server",
        url="https://my-mcp-server.example.com/mcp",
        timeout_seconds=120,
    )

Any MCP server that speaks the Model Context Protocol over HTTP or Streamable HTTP works with LumiBot. There are over 20,000 MCP servers available today covering news, economic data, filings, social sentiment, and more.

Built-in Tools
--------------

LumiBot includes a full set of built-in trading tools that are available to every agent **by default**. You do not need to list them explicitly. Even when you add custom tools via ``@agent_tool`` or MCP servers, the built-in tools remain available.

The built-in tools cover everything a trading agent needs:

- **Account:** ``account.positions``, ``account.portfolio`` -- current holdings and portfolio state
- **Market data:** ``market.last_price``, ``market.load_history_table`` -- real-time quotes and historical bars
- **DuckDB:** ``duckdb.query`` -- SQL queries over time-series data loaded into DuckDB tables
- **Orders:** ``orders.submit``, ``orders.cancel``, ``orders.modify``, ``orders.open_orders`` -- full order management
- **Documentation:** ``docs.search`` -- search LumiBot's own API docs for guidance

These tools give the agent access to positions, prices, history, and order execution without any setup. If you want to add external data on top of these, use ``@agent_tool`` or add MCP servers.

System Prompts
--------------

LumiBot handles all the common instructions internally through its base prompt. The base prompt tells the agent:

- Whether the run is a backtest or live trading
- The current datetime and timezone
- Current positions, cash, and portfolio values
- Rules about look-ahead bias and backtesting safety
- Default investor policy (conviction over activity, no overtrading)
- Position sizing, order execution, and limit order preferences
- DuckDB conventions and tool usage guidance

**Your system prompt should be 2-3 sentences about your strategy.** LumiBot handles the rest.

.. code-block:: python

    system_prompt=(
        "Use economic data to decide whether capital should be in TQQQ "
        "or a defensive asset like SHV. Check interest rates, inflation, "
        "and growth conditions. This is a binary allocator."
    )

Do not repeat instructions about position sizing, time safety, or tool usage. LumiBot already covers those in the base prompt.

Agent Handoffs
--------------

Multi-agent strategies often pass one agent's output into the next agent. For
example, an evidence researcher may hand a research pack to a bull researcher,
then a bear researcher, then a portfolio manager. These handoffs should be
large enough to preserve useful evidence while still being concise enough for
the next model call.

Prefer prompt instructions and structured output requests:

.. code-block:: python

    result = self.agents["evidence_researcher"].run(
        task_prompt=(
            "Build a structured evidence handoff. "
            "Keep it under context.handoff_target_tokens tokens. "
            "Do not pad the answer just to fill the budget."
        ),
        context={"handoff_target_tokens": 24000},
    )

    evidence_pack = result.summary or result.text

``handoff_target_tokens`` is the prompt target. It does not force the model to
use that many tokens. It tells the model the upper bound for a complete,
structured handoff. A good model can still return 5,000 or 8,000 tokens when
that is enough.

Do not silently truncate handoffs or tool results in order to make a backtest
fit a provider context window. Silent truncation changes the evidence the next
agent sees and can turn a trading-quality benchmark into a benchmark of the
truncation policy. If a handoff is too large, prefer narrower tools, better
role prompts, provider-appropriate model selection, or a clear failure with
diagnostics.

For 128K-context models, think about the combined context, not just one
handoff. If the portfolio manager receives evidence, bull, and bear handoffs,
three 32K-token handoffs can already consume roughly 96K tokens before the
system prompt, tool schemas, runtime context, and the portfolio manager's own
output.

Do not add hidden runtime tool-call budgets to trading benchmarks. Blocking
tools can invalidate results by preventing execution tools, such as order
submission, from running. If you need to control paid benchmark spend, use an
explicit outer run cap such as ``LUMIBOT_AGENT_MAX_MODEL_CALLS`` and treat the
run as failed when the cap is reached.

DuckDB and Time-Series Data
----------------------------

When the agent needs to analyze historical price data, LumiBot loads it into DuckDB tables automatically. The agent can then query these tables with SQL instead of reading raw bar data in the prompt.

This is handled by the base prompt and the built-in ``market.load_history_table`` and ``duckdb.query`` tools. The agent loads a price history table by symbol and timeframe, then queries it with standard SQL for moving averages, volatility, or any other analysis. You do not need to configure DuckDB -- it is part of the default agent runtime.

Replay Cache
------------

In backtesting mode, LumiBot caches every agent run. When a subsequent backtest hits the same combination of prompt, context, model, tools, and simulated timestamp, the cached result is returned instantly without calling the LLM or any external tool.

This means:

- **Deterministic backtests.** The same inputs always produce the same outputs.
- **Fast warm reruns.** A cached backtest that took 30 minutes on the first run can complete in seconds.
- **Cost control.** No duplicate LLM API calls or external API calls on repeated runs.

The replay cache is automatic. No configuration needed.

Observability
-------------

Every agent run produces a structured trace that records:

- The full prompt surface (base prompt + system prompt + context)
- Every tool call and tool result
- Any observability warnings (e.g., future-dated data in a backtest)
- The agent's summary and reasoning
- Cache hit/miss status
- DuckDB query metrics

A compact summary log line is emitted for every run. For deeper debugging, inspect the full JSON trace file. See :doc:`agents_observability` for the complete debugging workflow.

Canonical Demos
---------------

LumiBot ships short one-agent demos in ``lumibot/example_strategies/agent_*.py``. Each is a few sentences of plain English, about 30 lines, and uses only LumiBot's built-in tools:

1. **News Sentiment** (``agent_news_sentiment.py``) -- buys the well-known stocks with the strongest good news.
2. **Trend** (``agent_macro_risk.py``) -- holds TQQQ or SHV based on the price trend.
3. **Momentum and News** (``agent_momentum_allocator.py``) -- holds TQQQ or SHV based on trend and news.
4. **M2 Liquidity** (``agent_m2_liquidity.py``) -- holds TQQQ or SHV based on the Federal Reserve's money supply data.

See :doc:`agents_canonical_demos` for all of them.

The demo files are located at ``lumibot/example_strategies/agent_*.py`` and can be run directly after setting the required environment variables.

Frequently Asked Questions
--------------------------

**Can I backtest an AI trading agent?**

Yes. LumiBot lets an AI agent reason, call tools, and execute trades on every bar during a backtest. The agent runs inside ``on_trading_iteration()``, receives point-in-time market state, and uses tools to make decisions -- all within the backtest simulation. A built-in replay cache makes warm reruns deterministic and fast.

**What makes LumiBot different from other AI trading frameworks?**

Most alternatives either put the LLM outside the backtest loop (QuantConnect), have no backtesting at all (CrewAI, AutoGen, LangGraph), or are hobby scripts with no infrastructure. LumiBot runs the AI agent inside the backtest simulation on every bar, with ``@agent_tool`` for reliable external data, MCP server support, replay caching, DuckDB time-series queries, and full observability -- all with the same code for backtest and live.

**What AI models are supported?**

LumiBot ships with first-class support for Gemini, OpenAI (GPT), xAI (Grok), Anthropic (Claude), and any other provider covered by LiteLLM (~100 providers). You pick the model per agent via the ``default_model`` parameter when creating your agent.

The default is ``"openai/gpt-6-luna"`` with medium reasoning effort. Gemini ids (e.g. ``"gemini-3.5-flash-lite"``) take Google ADK's native fast path. Anything else is automatically routed through LiteLLM using the provider-prefixed id format:



- xAI Grok: ``"xai/grok-4.20-0309-reasoning"`` (Grok 4.2, reasoning on, 2M ctx), ``"xai/grok-4-1-fast-reasoning-latest"`` (cheap/fast), or ``"xai/grok-4-latest"`` (older) -- requires ``XAI_API_KEY`` or ``GROK_API_KEY``


The replay cache keys on the model id, so swapping providers on the same backtest produces fresh runs rather than stale cross-model replays. Tool calling is normalized across providers by LiteLLM, so your ``@agent_tool`` functions work unchanged regardless of which model you pick.

**How do I get started?**

Install LumiBot, set ``OPENAI_API_KEY`` in your environment, copy the Quick Start example on this page, and run it. The M2 Liquidity Strategy example is a complete, runnable strategy file. Provider-specific variants are available for OpenAI, Grok, and Anthropic. See :doc:`agents_quickstart` for additional patterns and :doc:`agents_canonical_demos` for the reference demo strategies.

**What API keys do I need?**

At minimum, one model provider key matching the ``default_model`` you set: ``OPENAI_API_KEY`` for GPT models (the default), ``GEMINI_API_KEY`` for Gemini, ``XAI_API_KEY`` or ``GROK_API_KEY`` for Grok, or ``ANTHROPIC_API_KEY`` for Claude. If your ``@agent_tool`` functions call external APIs, you also need those keys -- for example ``ALPACA_API_KEY`` and ``ALPACA_API_SECRET`` for Alpaca data APIs. Macro-data examples and built-in FRED tools require ``FRED_API_KEY`` so LumiBot can use the official FRED/ALFRED API and request point-in-time vintage observations in backtests.

**How do I set up my environment?**

Create a ``.env`` file in your project directory with your API keys (e.g., ``OPENAI_API_KEY=your_key_here``). LumiBot reads environment variables at startup. You can also export them in your shell. For backtesting, set ``BACKTESTING_DATA_SOURCE`` in ``.env`` or use ``datasource_class=None`` to defer to the environment configuration.

**Can I use this for live trading?**

The same ``Strategy`` class can be used in a backtest and with a supported
broker (Alpaca, Interactive Brokers, Tradier, Schwab, and others). The startup
code must select the path: ``Strategy.backtest(...)`` for history, or construct
the strategy with a broker and call ``run_live()`` or ``Trader.run_all()`` for
broker execution. Some example files include only a backtest runner. See
:doc:`strategy_run_modes` before running one directly.

**Does it work with my broker?**

LumiBot supports Alpaca, Interactive Brokers, Tradier, Schwab, Tradovate, TopstepX futures (via ProjectX), Bitunix, and selected CCXT crypto paths. Coinbase, Kraken, and WEEX have auto-detected credential paths; KuCoin, Binance, and BitMEX have documented manual CCXT setup paths; Kraken, Binance, KuCoin, BitMEX, Bybit, and OKX have documented backtesting examples. Lumibot does not claim support for every CCXT exchange. Any broker supported by LumiBot works with AI agents. The agent submits orders through the standard LumiBot order execution pipeline.

**What is @agent_tool?**

``@agent_tool`` is a decorator that wraps a Python method as a callable tool the AI agent can invoke during its reasoning loop. You provide a name and description, write a standard method with type hints and a docstring, and the decorator handles the rest. The function's source code is automatically included in the tool description so the AI can see parameters, defaults, and implementation details.

**How does the agent know what parameters my tool accepts?**

``@agent_tool`` automatically includes the function's entire source code in the tool description sent to the AI model. The AI sees your type hints, default values, and docstring. Write a clear docstring with a Google-style ``Args`` section and the AI will understand how to call your tool.

**Do I need to list built-in tools?**

No. All built-in tools (positions, portfolio, prices, orders, DuckDB, docs) are always included automatically. When you pass custom tools via ``tools=[self.my_tool]``, they are added alongside the built-in tools. You only need to list your custom ``@agent_tool`` functions. Outbound web and browser tools are the exception: they need ``allow_network=True``.

**Can I use multiple custom tools?**

Yes. Pass a list of tools when creating the agent: ``tools=[self.tool_a, self.tool_b, self.tool_c]``. The Macro Risk and Momentum Allocator demos both use multiple ``@agent_tool`` functions in a single strategy. There is no hard limit on the number of custom tools.

**What REST APIs can I wrap with @agent_tool?**

Any REST API that returns JSON or text. The canonical demos wrap Alpaca News API, Alpaca Bars API, Alpaca Screener API, and other HTTP services. For FRED macro data, prefer Lumibot's built-in FRED tools because they use the official API with realtime vintage parameters for point-in-time backtests. You can wrap Alpha Vantage, your own internal services, SEC EDGAR, social sentiment APIs, or anything else accessible over HTTP.

**How do I add authentication to my tool?**

Read API keys from environment variables inside your ``@agent_tool`` function using ``os.environ.get("MY_API_KEY")``. Pass them as headers or query parameters in your ``requests`` call. See the Alpaca demos for examples that use ``APCA-API-KEY-ID`` and ``APCA-API-SECRET-KEY`` headers.

**What happens if my tool returns an error?**

Return a dictionary with an ``"error"`` key (e.g., ``return {"error": str(e)}``). The agent sees the error and can decide to retry, try a different approach, or proceed without that data. An observability warning is also recorded in the trace. Wrap your HTTP call in a try/except block to handle network failures gracefully.

**Can I use MCP servers instead of @agent_tool?**

Yes. Pass an ``MCPServer`` object with a URL when creating the agent. However, ``@agent_tool`` is the recommended primary pattern because you control the HTTP call directly, it works reliably in both backtests and live trading, and it does not require external server infrastructure.

**What is the difference between @agent_tool and MCP servers?**

``@agent_tool`` wraps a Python method that makes HTTP calls via ``requests`` -- you control the code, it runs in-process, and it works reliably in backtests. MCP servers are external services that speak the Model Context Protocol over HTTP. MCP servers are useful when a third party provides a dedicated server or you need access to one of the 20,000+ public MCP servers, but ``@agent_tool`` is more reliable for backtesting and gives you full control.

**How long should my system prompt be?**

Two to three sentences describing your strategy intent. For example: what data to use, what assets to trade, and what the allocation logic should be. LumiBot handles position sizing, DuckDB guidance, backtesting safety, time-awareness, and the default investor policy in its base prompt.

**What should I put in the system prompt?**

Describe your strategy's thesis and the assets it trades. Do not repeat instructions about position sizing, order execution, look-ahead bias, or tool usage -- LumiBot covers all of that in the base prompt. A good example: ``"Use economic data to decide between TQQQ and SHV. Check interest rates, inflation, and growth conditions."``

**What does LumiBot handle automatically in the base prompt?**

The base prompt tells the agent whether the run is a backtest or live, the current datetime and timezone, current positions and cash, rules about look-ahead bias, the default investor policy (conviction over activity, no overtrading), risk and drawdown discipline (risk-adjusted returns over raw returns, recovery math, cut losers, no chasing after drawdowns, Sharpe/Sortino/Calmar framing), position sizing and limit order preferences, and DuckDB conventions and tool usage guidance.

**Can I override the default investor policy?**

The base prompt includes a default policy favoring conviction over activity and discouraging overtrading. Your system prompt can direct the agent toward different behavior -- for example, telling it to rebalance daily or trade more aggressively. The system prompt is added on top of the base prompt, so your instructions take priority for strategy-specific guidance.

**How do I make the agent more aggressive or more conservative?**

Add explicit direction in your system prompt. For a more aggressive agent: ``"Trade actively. Rebalance into high-conviction positions quickly."`` For a more conservative agent: ``"Only trade when evidence is overwhelming. Prefer holding cash or SHV when uncertain."`` The agent follows your prompt guidance.

**How does backtesting work with AI agents?**

The agent runs inside ``on_trading_iteration()`` on every bar (e.g., every trading day if ``sleeptime="1D"``). On each bar, the agent receives point-in-time market state, calls tools (both built-in and custom), reasons over the data, and submits orders. The backtest simulation processes those orders at simulated market prices. The replay cache makes warm reruns deterministic.

**How does the agent avoid looking into the future during backtests?**

LumiBot injects the simulated datetime into the agent's context and the base prompt includes explicit rules about look-ahead bias. The observability system also flags future-dated data warnings if a tool result references data published after the simulated backtest time. Your ``@agent_tool`` functions should respect date parameters to avoid requesting future data.

**What is the replay cache?**

In backtesting mode, LumiBot caches every agent run keyed by a SHA-256 hash of the prompt, context, model, tool surface, and simulated timestamp. When a subsequent backtest hits the same combination, the cached result is returned instantly without calling the LLM or any external tool. This makes warm reruns deterministic, fast, and cost-free.

**How do I clear the cache for a fresh run?**

Delete the replay cache directory. On macOS the default location is ``~/Library/Caches/lumibot/1.0/agent_runtime/replay/``. Set ``LUMIBOT_CACHE_FOLDER`` before importing lumibot to store caches somewhere else. After clearing, the next run will make fresh LLM and tool calls.

**How long does a backtest take?**

A cold run (no cache) depends on the number of bars, the number of tool calls per bar, and the LLM response time. A six-year daily backtest with one tool call per bar might take 20-40 minutes on the first run. A warm run (fully cached) completes the same backtest in seconds because no LLM or external API calls are made.

**Can I speed up backtests?**

Use the replay cache -- after the first cold run, all subsequent runs with the same inputs are near-instant. You can also reduce the date range, increase the ``sleeptime`` to trade less frequently, or use a faster model. Keeping your ``@agent_tool`` functions fast (short timeouts, efficient parsing) also helps.

**What data sources work for backtesting?**

Set ``datasource_class=None`` to use the data source from your ``.env`` file (via ``BACKTESTING_DATA_SOURCE``). For standalone examples, use ``YahooDataBacktesting``. LumiBot also supports ThetaData, Polygon, and other data sources for backtesting. The data source controls price bars and market data; your ``@agent_tool`` functions provide any additional external data.

**How do I see what the agent is doing?**

Every agent run emits a compact summary log line with the agent name, model, cache status, tool call count, warning count, and the agent's summary conclusion. The queryable record is ``*_agent_detail.parquet``. The ``call_summary`` row includes ``effective_system_prompt``. Raising ``LUMIBOT_LOG_LEVEL`` only changes printed logs. See :doc:`agents_observability` for the full debugging workflow.

**What are agent traces?**

``*_agent_detail.parquet`` is one table for the whole run: a ``call_summary`` row per AI call, plus rows for thinking, text, tool calls, and tool results. A JSON trace is also written per call. Both record the prompt (``effective_system_prompt``), every tool call and result, the summary, warnings, and cache status.

**Where are trace files stored?**

A backtest writes ``*_agent_detail.parquet`` next to the tear sheet. Live and paper files are under ``~/Library/Caches/lumibot/1.0/agent_runtime/`` on macOS. Set ``LUMIBOT_CACHE_FOLDER`` before importing lumibot to move that folder. The per-call JSON path is ``(result.payload or {}).get("trace_path")``. Summaries are also written to ``agent_run_summaries.jsonl``. ``LUMIBOT_LOG_LEVEL`` does not change either location.

**How do I debug a bad trade?**

Open ``*_agent_detail.parquet`` for that run and read ``effective_system_prompt``, the tool rows, and the summary. Look for warnings (future-dated data, no tools called, unsupported orders). Compare the summary to the trade. See :doc:`agents_observability`. Raising ``LUMIBOT_LOG_LEVEL`` will not add this record.

**Why is my agent not trading?**

Check the agent's summary in the logs -- it may have decided not to trade because conviction was low. The default investor policy in the base prompt encourages conviction over activity. If you want more frequent trading, adjust your system prompt to be more directive. Also verify that your tools are returning valid data by inspecting the trace.

**Why is my agent only buying SHV?**

SHV is a common defensive parking asset used in the demo strategies. If the agent only buys SHV, it means the agent is not finding enough conviction to take risk. Check whether your tool is returning useful data (inspect the trace), whether the system prompt is clear about when to be risk-on, and whether the market data covers the right date range.

**How much does it cost to run?**

Cost depends on the LLM provider and model, the number of bars in your backtest, and how many tool calls the agent makes per bar. The first cold run of a long backtest makes one or more model calls per bar, so check your provider's current pricing and start with a short date range. Warm reruns cost nothing because the replay cache eliminates all LLM and external API calls.

**How can I reduce API costs?**

Use the replay cache -- compatible cached decisions avoid another model call. Use cost-effective models (e.g., ``openai/gpt-6-luna``). Keep your backtest date range focused during development. Reduce the number of tool calls by making your tools return comprehensive data in a single call rather than requiring multiple round trips.

**How does replay caching reduce costs?**

The replay cache stores every agent run result keyed by a hash of the inputs. When the same prompt, context, tools, model, and timestamp appear again, the cached result is returned with zero LLM calls, zero external API calls, and zero cost. A cold backtest that costs a few dollars becomes free on every subsequent warm run.

Error Handling and Reliability
------------------------------

.. note::

   This section describes error handling **specific to AI agent calls**. The rest of LumiBot's main-loop error handling (strategy executor, brokers, data sources) is unchanged. The behavior below is scoped to ``AgentHandle.run()`` and ``GoogleADKRuntime.run()``; it does not alter how non-agent code paths react to exceptions.

LumiBot's AI agent stack has four timeout/retry/safety layers that together keep live trading alive through provider outages and surface backtest-time bugs clearly:

1. **Provider request timeout.** Each individual model request has a default **10 minute** timeout. Native Gemini models receive this as ``google.genai.types.HttpOptions(timeout=...)``. LiteLLM-backed providers receive it as LiteLLM's ``timeout`` argument. This prevents one wedged provider call from freezing an agent for the full run budget.

2. **LiteLLM-level HTTP retries.** When using non-Gemini providers, LiteLLM retries each individual HTTP call 3 times with provider-aware backoff (429 Retry-After awareness, capped exponential). Configured automatically in ``_configure_litellm_quietly`` (``num_retries=3``, ``drop_params=True``, ``suppress_debug_info=True``).

3. **Runtime-level attempt retries.** ``GoogleADKRuntime.run()`` retries the full agent call up to **10 times** live (2 in backtests, 1 for agents with order tools) with capped exponential backoff (2s, 3s, 5s, 10s, 20s, 30s, 45s, 60s, 60s, 60s). This covers session-setup errors, ADK runner glitches, and provider 5xx storms that LiteLLM's inner retry couldn't fix. Only transient and unknown errors retry; auth/config/billing errors surface immediately so we do not waste 5 minutes retrying a wrong API key.

   **Rate limits (HTTP 429)** get their own bounded budget of **6 attempts** in every mode, including backtests and trading agents. Each wait honors the provider's ``Retry-After`` header (or "try again in Ns" in the message), capped at 120 seconds per wait. A run is never retried after it already submitted, changed or cancelled an order, because a retry could duplicate the order; the next bar re-evaluates instead. An explicit ``LUMIBOT_AGENT_MAX_RUN_ATTEMPTS`` still caps everything.

4. **Strategy-level safety net with live-vs-backtest branch.** ``AgentHandle.run()`` wraps the runtime call in a final catch. Behavior depends on two things: the error category and whether the strategy is in backtest mode or live.

Timeout configuration
~~~~~~~~~~~~~~~~~~~~~

The provider request timeout is different from the full agent run timeout:

- ``model_request_timeout_seconds`` controls one model/API request. Default: ``600`` seconds.
- ``run_timeout_seconds`` controls the whole agent run, including model calls, tool calls, and retries. Default: ``1800`` seconds.

Set these when creating an agent:

.. code-block:: python

    self.agents.create(
        name="researcher",
        model="openai/gpt-6-luna",
        system_prompt="Research the best trade.",
        model_request_timeout_seconds=600,
        run_timeout_seconds=1800,
    )

Or override them for one call:

.. code-block:: python

    self.agents["researcher"].run(
        task_prompt="Run a deeper research pass.",
        model_request_timeout_seconds=900,
        run_timeout_seconds=2400,
    )

Output token limit
~~~~~~~~~~~~~~~~~~

By default LumiBot sends no output length, so the model answers as long as it needs. Anthropic models require one, so they get the model's real output limit. ``max_output_tokens`` lets a strategy set a limit anyway; it is capped at the model's real output limit (for example 16,384 for ``openai/gpt-4o``). Set it on ``create()`` for every run, or on ``run()`` for one call:

.. code-block:: python

    self.agents.create(name="analyst", model="openai/gpt-6-luna", max_output_tokens=4000)
    self.agents["analyst"].run(task_prompt="One-line verdict.", max_output_tokens=800)

Reasoning models count their thinking tokens against this limit, so keep it generous for deep research agents.

Advanced operators can also set ``LUMIBOT_AGENT_MODEL_REQUEST_TIMEOUT_SECONDS`` and ``LUMIBOT_AGENT_RUN_TIMEOUT_SECONDS``. A non-positive value disables that timeout. LumiBot logs every cold agent call with the effective timeout values and logs the latency to the first ADK event, which helps distinguish a stuck provider request from an agent that is actively calling tools.

Error classifier buckets
~~~~~~~~~~~~~~~~~~~~~~~~

Every exception from an agent call is classified by ``_classify_agent_error`` into one of five buckets:

- ``auth`` -- missing or invalid API key, permission denied (401, 403)
- ``config`` -- bad model id, malformed prompt, context-window exceeded, invalid payload (400, 404, 422)
- ``billing`` -- out of credits, payment required, quota exhausted (402, 429 with ``insufficient_quota``, 403 with billing/credits keywords)
- ``transient`` -- 5xx, rate-limit bursts, timeouts, connection errors
- ``unknown`` -- anything not matched; treated as transient (safe default)

The classifier looks at the exception class name, HTTP status code (if the provider SDK attached one), and message substring keywords (``insufficient_quota``, ``credits``, ``billing``, ``payment``, ``no credits``) so that a 403 returned with a billing message is correctly classified as ``billing`` rather than ``auth``.

Backtest vs. live behavior
~~~~~~~~~~~~~~~~~~~~~~~~~~

+--------------+------------------------------------------+----------------------------------+
| Category     | Backtest                                 | Live                             |
+==============+==========================================+==================================+
| ``auth``     | **Crash loud** with env-var guidance.    | Log + skip iteration.            |
+--------------+------------------------------------------+----------------------------------+
| ``config``   | **Crash loud** with model/prompt hint.   | Log + skip iteration.            |
+--------------+------------------------------------------+----------------------------------+
| ``billing``  | **Crash loud** with provider billing URL.| Log + skip iteration.            |
+--------------+------------------------------------------+----------------------------------+
| ``transient``| Log + skip iteration, counted and shown  | Log + skip iteration.            |
|              | in ``agent_health`` (see below).         |                                  |
+--------------+------------------------------------------+----------------------------------+
| ``unknown``  | Log + skip iteration, counted (safe      | Log + skip iteration.            |
|              | default).                                |                                  |
+--------------+------------------------------------------+----------------------------------+

**Live trading invariant**: an AI agent call never stops a live trading bot. Ever. Even a completely missing API key will log an error and continue — the operator can fix the env var and the bot resumes on the next iteration without a process restart. This is intentional: shutting down a live bot with real money at risk because of a provider hiccup is unacceptable.

**Backtest philosophy**: surface bugs loudly. A silent +0% tearsheet caused by a wrong API key is worse than a clear error message — the user just started the run, can fix it, and re-run. Transient errors (including rate limits that outlast their retries) skip the bar, but never silently: a backtest with skipped bars ran with missing AI decisions, and its results must say so.

Skipped bars in backtest results
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Every skipped agent call is counted and reported in three places:

- the log, as ``BACKTEST INCOMPLETE: agent '<name>' ... could not run at <time>``;
- the tear sheet's **Parameters Used** panel, as ``agent_<name>_skipped_runs`` and ``agent_skipped_runs_total``;
- the backtest ``settings.json``, as an ``agent_health`` block:

.. code-block:: json

    {
      "agent_health": {
        "complete": false,
        "skipped_runs": 3,
        "skipped_runs_by_agent": {"analyst": 3},
        "by_category": {"transient": 3},
        "skipped": [{"agent": "analyst", "datetime": "2026-09-23T10:00:00-04:00", "category": "transient", "error_class": "RateLimitError"}],
        "model_calls": 120
      }
    }

``complete: true`` means every agent call in the backtest produced a decision.

Skipped iteration result shape
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

When the safety net returns a graceful skip, the ``AgentRunResult`` includes:

- ``summary`` starting with ``"RESULT: Skipped this iteration. Agent call failed (category=...)"``
- a text event with payload ``{"runtime_error": True, "error_category": "...", "error_class": "...", "error_message": "...", "traceback": "..."}``
- a warning in ``result.warnings`` with ``kind="agent_runtime_failure_skipped"`` and the category
- ``cache_key = None`` (failures are never cached — next iteration retries fresh)

Strategy authors can count skipped iterations with ``len([w for w in result.warnings if w.get("kind") == "agent_runtime_failure_skipped"])`` for post-run analysis.

Model id visibility in tearsheets
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Every time you call ``self.agents.create(name=..., default_model=...)``, the framework auto-populates ``self.parameters[f"agent_{name}_model"]`` with the resolved model id. This shows up automatically in the tearsheet's **Parameters Used** panel so every AI backtest self-identifies which model produced which tearsheet. Multi-agent strategies get one key per agent.

Token usage and audit trail
~~~~~~~~~~~~~~~~~~~~~~~~~~~

LumiBot also writes AI usage details for every agent run during a backtest:

- The tearsheet's **Parameters Used** panel shows running totals for each agent:

  - ``agent_<name>_calls``
  - ``agent_<name>_input_tokens``
  - ``agent_<name>_output_tokens``
  - ``agent_<name>_total_tokens``
  - ``agent_<name>_thinking_tokens``
  - ``agent_<name>_cached_input_tokens``
  - ``agent_<name>_uncached_input_tokens``
  - ``agent_<name>_latency_ms_avg``
  - ``agent_<name>_tool_calls``
  - ``agent_<name>_cache_hits``
  - ``agent_<name>_detail_parquet``

- A single detailed tabular artifact is written beside the normal backtest artifacts using the same base filename pattern:

  - ``<run>_agent_detail.parquet``

The Parquet file is the canonical machine-readable audit artifact used by BotSpot/MCP query tooling. LumiBot does not estimate provider pricing in this file because model prices change; it records raw token usage only.

Each agent call gets one ``call_summary`` row plus one row per model event inside the call. This avoids repeating call-level token totals on every tool row while still preserving the event timeline. The file includes:

- prompt/context fields (user system prompt, effective prompt, task prompt, runtime context)
- event kind (``call_summary``, ``thinking``, ``text``, ``tool_call``, ``tool_result``, ``usage`` when present)
- event text
- tool name
- flattened tool/event details in normal columns
- full event payload JSON for exact forensic inspection
- input/output/total token counts on the ``call_summary`` row
- cached/uncached input token counts when the provider reports them
- thinking token counts when the provider exposes them
- latency fields for the full call and first model event
- cache-hit flag and warnings

This file is meant to answer practical debugging questions after a backtest:

- What exactly did the agent say?
- What tools did it call?
- What came back from those tools?
- How many tokens did that call use?
- How many input tokens were cached vs. uncached?
- How long did the call take?

Thinking text is captured when the provider/SDK exposes it. Gemini thought summaries are requested automatically. Other providers may expose only thinking token counts and not the actual thought text.

Provider prompt caching
~~~~~~~~~~~~~~~~~~~~~~~

LumiBot has two separate cache layers:

- The LumiBot replay cache skips the entire agent call on identical warm backtests.
- Provider prompt caching reduces cost/latency during cold backtests and live trading when the static prompt prefix repeats.

The agent runtime keeps the large, stable instructions and tool definitions at the beginning of the request and moves dynamic fields such as current datetime, runtime mode, positions, orders, memory, task prompt, and user context into later request sections. This improves provider prefix-cache hit rates without changing strategy behavior.

Provider-specific routing:

- OpenAI models receive a stable ``prompt_cache_key`` plus ``prompt_cache_retention="24h"`` through LiteLLM.
- xAI/Grok models receive a stable ``x-grok-conv-id`` header through LiteLLM.
- Gemini native models use Gemini's implicit caching path. Explicit ADK context caching is a future optimization; the runtime already records Gemini ``cached_content_token_count`` when the provider reports it.

Use ``scripts/run_agent_prompt_cache_probe.py`` to verify provider-reported cache behavior with real calls:

.. code-block:: bash

    python scripts/run_agent_prompt_cache_probe.py --model openai/gpt-6-luna
    # Optional: compare another model
    python scripts/run_agent_prompt_cache_probe.py --model openai/gpt-5.4-mini

The probe bypasses LumiBot's replay cache, sends repeated calls with the same long static prefix, and prints input tokens, cached input tokens, uncached input tokens, output tokens, and latency for each call.

Built-in Alpaca news tool
~~~~~~~~~~~~~~~~~~~~~~~~~

Strategies can use ``BuiltinTools.news.alpaca_news()`` to give an agent access to the Alpaca News API without writing a custom wrapper:

.. code-block:: python

    from lumibot.components.agents import BuiltinTools

    self.agents.create(
        name="trader",
        system_prompt=(
            "Use Alpaca news and market tools to make trading decisions. "
            "Scan headlines and summaries first. If a story matters, fetch full article content before trading."
        ),
        tools=[BuiltinTools.news.alpaca_news()],
    )

The tool uses the active Alpaca broker credentials when the strategy is running on Alpaca, including OAuth connections. If the active broker is not Alpaca, set bring-your-own-key news credentials with ``ALPACA_NEWS_API_KEY`` and ``ALPACA_NEWS_API_SECRET``. If neither path is available, LumiBot logs a warning and does not expose ``alpaca_news`` to agents. It defaults ``end`` to the current simulated datetime in backtests and clamps future ``end`` values to avoid look-ahead. The response includes ``requested_end``, ``effective_end``, and ``lookahead_clamped`` so you can audit the exact window used.

Alpaca news is historical symbol/date-window retrieval, not keyword search. The API supports ``symbols``, ``start``, ``end``, ``limit`` (max 50), ``sort``, ``include_content``, ``exclude_contentless``, and ``page_token``. For broad market context, query market ETF proxies such as ``SPY,QQQ,DIA,IWM``; for sector context, query sector ETFs such as ``XLK,SMH`` (tech/semis), ``XLF,KRE`` (financials/banks), ``XLE,USO`` (energy), ``XLV,XBI`` (healthcare/biotech), ``TLT,IEF,SHY`` (rates/bonds), or ``GLD,SLV,DBC`` (gold/commodities).

Use a two-step workflow:

1. Scan with ``include_content=False``. Use ``limit=10`` to ``20`` for focused single-symbol checks and ``limit=30`` to ``50`` for broad market or sector scans. This returns headlines, summaries, URLs, sources, timestamps, symbols, and ``next_page_token`` without dumping long article bodies into the model context.
2. If a story looks important, call again for the same or narrower window with ``include_content=True`` and usually ``exclude_contentless=True``. Full article content is returned without truncation unless you explicitly pass ``content_max_chars``.
3. If ``next_page_token`` is present and the first page does not provide enough evidence, call again with ``page_token=next_page_token``.

Do not trade from one weak or noisy article. News can be sparse for single stocks, so broaden from the stock to its sector or market ETF when needed, compare article timestamps against the simulated datetime, and use ``page_token`` when the first page does not provide enough evidence.

Complete runnable example. The prompt is plain English; the agent finds and uses the news tool on its own:

.. literalinclude:: ../lumibot/example_strategies/agent_alpaca_news_builtin.py
   :language: python



To run the live proof that validates historical relevance, full-content retrieval, and the resulting ``*_agent_detail.parquet`` artifact:

.. code-block:: bash

    python scripts/run_alpaca_news_ai_proof.py --model openai/gpt-6-luna

Shown in full with attribution under the source's licence. Licence: GPL-3.0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.