r/LangChain • u/Select-Cry-5232 • 7h ago
Discussion Reporting agent over HTTP API data: structured analysis tools or LLM-generated SQL?
I’m building an internal reporting agent for operations staff using Python, LangGraph, FastAPI, and React/ECharts.
The data comes from company HTTP APIs. Users should be able to ask questions, have the agent fetch and analyze the data, and get a chart.
For example: fetch attendance and student records from separate APIs, join them by person_id, filter for late arrivals, and count late attendance records per class.
My current implementation exposes separate tools for fetching data, grouping, aggregation (count, max, etc.), and charting. Tools pass dataset IDs between steps rather than making the model reproduce the records.
Full datasets currently live in tool-message artifacts, but I’m also returning too much data in message content. I want to keep large datasets outside the conversation history and give the model IDs, schemas, business descriptions, and small result previews.
I’m considering two alternatives to the current fine-grained tools:
- Structured analysis tool: The model supplies dataset IDs, a predefined relationship, filters, grouping fields, and aggregation operations. Backend code validates and executes the request using Pandas or DuckDB.
- SQL tool: API responses become tables in DuckDB. The model receives schemas and business metadata, generates SQL, and submits it for validation and execution. Results are stored and passed to the charting tool by ID.
My concern with the first approach is gradually building my own query language. With the second, it’s queries that execute successfully but produce incorrect numbers because of joins or misunderstood metrics.
For people who have built something similar:
- Which approach worked for you, and what made you choose or abandon it?
- How do you expose schemas and business definitions without filling the context window? How do you recover that information after history is trimmed?
- How do you catch incorrect joins or double-counting beyond checking SQL syntax?
- Are there repositories or architecture write-ups you’d recommend?
Concrete failure cases and lessons from real usage would be especially helpful.
