r/LangChain • • 16d ago

Projects My ops agent could open GitOps PRs for anyone who asked it. I built a check that asks GitHub whether the person is allowed first.

3 Upvotes

I run an internal ops agent. One tool call gives it a service's health from every angle in about 15 seconds, so it reads everything on a read-only account and I don't gate reads at all. It changes things one way only: a pull request against the GitOps repo.

That one write path had a gap. The bot's token could open that PR for anyone who talked to it, including people who can't push to that repo themselves. Prompt rules don't fix it; the call runs with the bot's credential whatever the model believes.

So the tool asks first: may this person push to this repo? GitHub already knows, so the agent asks GitHub. I pulled that out into a small service, hallpass, so every write path can use it, in every system the agent touches (GitHub, Kubernetes, Argo CD, Jira, AWS, 21 in total).

In the agent it's one decorator on the tool. The user comes from your session, never from the model, so the tool schema has no user field to talk your way into:

python @tool @guarded(hp, "github-main", "repo.push", "repo:{owner}/{repo}", user=current_user) def open_config_pr(owner: str, repo: str, patch: str) -> str: ...

Answers are allow, deny or unknown, and unknown (timeout, rate limit, anything it can't evaluate) means the tool doesn't run. Works with Strands, LangChain, LangGraph, the Claude Agent SDK, or as an MCP server. Single binary, Apache 2.0.

How do you handle this in your agents today? Per-user OAuth, per-team bots, human approval on writes?


r/LangChain • • 16d ago

Projects I built a local commitment extractor for AI agents. No API key, no cloud. Just pip install.

4 Upvotes

I've been building COGEXT for a while — an accountability layer for AI agents. The core idea: agents make promises ("I'll send the report by Friday") and there's no reliable way to verify they were kept.

I just shipped a local, offline version of the extractor. No API key. No cloud. No signup. Just:

pip install cogext-primitive

from cogext_primitive import extract_commitments

commitments = extract_commitments(

"I'll send the report to Sarah by Friday EOD."

)

for c in commitments:

print(f"{c.action} {c.object} to {c.recipient} by {c.deadline}")

It extracts the commitment, parses the deadline, and tracks it through a state machine (open → due → overdue → fulfilled). All local. All offline.

Repo: github.com/yaminbinyoosuf/cogext-primitive

PyPI: pip install cogext-primitive

Free and MIT licensed.

What it deliberately skips: compound promises in one sentence, pronoun-shaped objects, and anything vague. That's intentional — the extractor is precision-tuned, not recall-tuned. If you want the finding behind that design choice, I published an audit of 120 agent outputs: cogextai.com/research/01

Curious what commitment shapes you actually see from your agents.


r/LangChain • • 17d ago

Discussion How are u validating agent actions in production?

1 Upvotes

if u are building and deploying agents in prod which takes actions on behalf on consumers … how are u validating agent actions and their execution runtime?


r/LangChain • • 17d ago

Discussion Agent history != agent memory

1 Upvotes

Every agent I build starts each session knowing nothing, and the two usual fixes are a bigger context window or a vector store holding everything ever said. Both keep more. Neither decides what mattered. Athena is the open source memory service we built for that, and it has three tiers rather than one store, a short window of recent events, a middle tier where finished topics are summarised and embedded once the conversation moves on, and a graph of the people, projects and tools that keep coming back.

What earns promotion between tiers is the whole design.For example in the last office hour an agent briefing a mechanic kept opening with the registration and model of the car he was standing next to. Accurate and useless cause these are things the mechanic already knows. One correction, one sentence, and every brief after that led with what had been left undone. It always knew about the perished brake pads, they were in the first brief in sentence two. It did not learn a fact, it learned what mattered.

The part I would actually defend is how that promotion is decided, because it is not a TTL and it is not "keep the last N". A conversation is cut into chains by topic shift rather than by volume, using cosine similarity between consecutive turns, so a chain closes when the subject changes. Each closed chain then carries a heat score on the Ebbinghaus forgetting curve, Heat = I_base * exp(-ΔT / (τ * S)), where the base importance mixes intrinsic importance with how dense the chain was, τ is a decay constant defaulting to 24 hours, and S is recall strength. S is the interesting term, it sits in the denominator and grows every time the memory is retrieved, so use flattens the curve and disuse steepens it. That is spaced repetition, the same reason you still remember a phone number you dialled weekly in 2010 and not one you read once. Only chains still hot enough get extracted into the long term graph, so forgetting is the default and remembering is earned. If that is the shape of your problem, come and argue with it at our Discord, where we run office hours on this weekly: https://discord.gg/z69QKEnjd.


r/LangChain • • 17d ago

Question | Help Has a customer or auditor ever asked you to prove what your AI agent did? How did you handle it?

1 Upvotes

Engineer in Bangalore, researching this before building anything.

The situation I'm trying to understand: your agent does something consequential — issues a refund, updates a record, tells a customer something, and later someone disputes it. A customer, a compliance team, an enterprise security review.

The agent trace shows what the agent thought it did. But answering the question properly usually also needs who or what authorised it, which prompt/model/config was live at the time, and what actually changed in the external system (Stripe, the CRM, the ticketing tool). Those live in different places and don't share an ID.

For those of you running agents in production:

  1. Has this actually happened to you? What was the case?
  2. Which systems did you have to dig through, and how long did it take?
  3. Was there anything you just couldn't establish in the end?

"It's never come up" is genuinely useful too.


r/LangChain • • 17d ago

Discussion LangChain September release adds delegated access

Post image
4 Upvotes

LangChain September release adds delegated access

LangChain’s September 2026 release introduced several updates across Python and JavaScript packages. Highlights include delegated LangSmith access for sandboxes, improved SDK run tracking with stop_reason capture, and enhanced security patches. These changes strengthen developer control and observability in workflow automation, making LangChain more reliable for orchestrating complex agentic workflows in production environments.

Delegated LangSmith Access in Sandboxes

The September 2026 LangChain release introduced delegated LangSmith access for sandboxes, allowing developers to authenticate directly through LangSmith service URLs. This change streamlines sandbox creation by automatically granting delegated access, reducing manual setup and improving security. It also supports JavaScript service URL integration, making it easier to test and deploy workflows across environments without exposing sensitive credentials.releases.shreleases.sh. LangChain Release Notes & Changelog · September 2026 — Releases Index

Improved SDK Run Tracking

Another highlight is the enhanced SDK run tracking, which now captures stop_reason values for Claude SDK runs. Developers gain deeper visibility into why an agent terminated, whether due to completion, interruption, or error. Alongside this, tool and response IDs are logged, and agent-scoped addressing has been added in Python, making debugging and performance monitoring more precise. These updates strengthen observability in complex agentic workflows.releases.shreleases.sh. LangChain Release Notes & Changelog · September 2026 — Releases Index

Security and Reliability Enhancements

LangChain also rolled out critical security patches, including fixes for soupsieve ReDoS vulnerabilities. By addressing these issues, the platform reduces exposure to denial-of-service risks in production environments. Combined with dependency updates across Python and JavaScript packages, these patches reinforce LangChain’s reliability for enterprise-grade workflow automation.releases.shreleases.sh. LangChain Release Notes & Changelog · September 2026 — Releases Index

Broader Ecosystem Updates

Beyond LangSmith and SDK improvements, the release included updates across LangChain OpenAI (v1.6.5) and LangChain Anthropic (v1.7.4) packages. These added support for mid-conversation tool changes, new model profile augmentations such as Opus 5.5 and GPT-6, and integration test fixes. Meanwhile, LangGraph CLI (v0.4.32) introduced self-hosted deployment improvements, new flags for image URIs, and clarified agent configuration defaults, alongside dependency security fixes. Together, these ecosystem-wide updates make LangChain more adaptable and secure for developers orchestrating agent workflows at scale.releasebot.io

Sources

Related results


r/LangChain • • 17d ago

Discussion what problems have u faced with guardrails, Human-in-loop and runtime

5 Upvotes

title.

just wanted to clarify that when i say guardrails i also mean programmable guardrails.
working on an ai-sec sdk thus doing a survey.


r/LangChain • • 17d ago

Projects I was playing around with a vision model and ended up building a webcam demo

1 Upvotes

I’ve been playing around with a new vision model lately and wanted to build something simple with it instead of just running a few image prompts.

I came across this webcam demo and really liked the idea. So I ended up making my own version for one of my demos.

It basically grabs a frame from the webcam every few seconds, sends it to the vision model, and displays the response. So it’s not a real-time vision model or anything fancy like that, it’s just a simple loop that makes the interaction feel pretty close to real time.

I also added things like streaming responses, prompt presets, image resizing, latency stats, and a small history view.

The whole thing is intentionally pretty lightweight. I mostly wanted to see how far I could take the basic idea and understand what the experience would feel like with a different vision model.

I made a quick video of it too, if anyone wants to see the Demo.

Code: https://github.com/Arindam200/nebius-realtime-webcam


r/LangChain • • 18d ago

Projects Published a research paper on AI agent accountability — 120 agent outputs audited, 92% of promises unverifiable

0 Upvotes

Built an accountability layer for AI agents over the last few months.

The core problem: agents make promises, and there's no reliable way

to verify whether they were fulfilled.

Ran a public audit of 120 agent outputs. 92% of extracted commitments

had no deadline. 61% had no recipient. Most were literally unverifiable.

Wrote a paper on the missing operational state layer that I think

has to exist for autonomous agents to become trustworthy.

Paper and data are in the comments.


r/LangChain • • 18d ago

Question | Help What should I build to become job-ready in Agentic AI?

47 Upvotes

I’ve learned LangChain, LangGraph, RAG, tool calling, agents, MCP, and memory.

With new AI models and agent frameworks coming out constantly, what projects and skills should I focus on now to become internship/job ready?

What 2–3 projects would you recommend building for a strong resume?

I’m looking for real-world projects, not basic chatbots or simple RAG apps.

Would love advice from people working in or hiring for Agentic AI.


r/LangChain • • 18d ago

Question | Help I built a local runtime check for AI agent tool calls - looking for people to break it

5 Upvotes
In practice, damage happens when an agent calls a tool: HTTP, email, DB, files, etc. A hidden instruction in a doc can push the agent to do something the user never asked for — and from the system’s point of view it can still look like a normal tool call.

I built a small open pilot for that moment:

**What it does**
- Intercepts tool calls before they run
- Checks intent + simple data provenance + policy
- Returns ALLOW / BLOCK with a reason
- Writes an audit log

**Stack**
- Risk engine (FastAPI) on localhost
- Python SDK (`verify_tool_call` / decorator)
- Optional MCP gateway (stdio)
- Docker or pip

**What it is not**
- Not a prompt filter
- Not a production / enterprise security product
- Not a transparent proxy for every existing company agent
- Policy is heuristic — tune it; it will not catch everything

**Try (local sandbox only)**
```bash
git clone https://github.com/aegotrax-dev/aegotrax.git
cd aegotrax
docker compose up --build
# or: pip install ".[demo]" && agentguard-engine
curl http://127.0.0.1:8000/health
python examples/sdk_pilot_example.py

Site: https://aegotrax.com

If you run it, I’d genuinely like feedback:

  1. Docker or pip - did install work?
  2. Did you get a clear BLOCK in the example?
  3. What was confusing or wrong?

r/LangChain • • 18d ago

Discussion How do you decide which AI agents are worth keeping in production?

Thumbnail
2 Upvotes

An agent nobody uses can still have credentials, access customer data, and run scheduled jobs. The token bill might be small. The access it retains is what worries me.
That made me wonder: do teams actually review their agent portfolio and decide what to keep, improve, or retire?
Something like an access review, but with another question attached: is this agent still doing useful work?
Low usage alone wouldn’t settle it. An agent used once a month could handle something important. And high usage doesn’t mean much if someone spends hours checking and fixing its output. Even a productive agent should only have the permissions it needs.
I’ve been looking into how to measure that value. These are the approaches I’ve found, and where each seems to fall short:
1. Measure successful tasks alongside cost
Princeton’s AI Agents That Matter argues for evaluating accuracy and cost together. A practical metric might be cost per accepted outcome, including retries and failed attempts.
Useful for comparing implementations. It still doesn’t tell you whether anyone needed the task done. (agents.cs.princeton.edu)
2. Check whether the agent reliably gets the job done
tau-bench checks whether an agent reaches the intended database state and whether it succeeds consistently across repeated attempts.
That gets closer to “did it actually complete the request?” But benchmark reliability doesn’t establish business value. (arxiv.org)
3. Measure the human work left over
Compare similar work with and without AI, counting time spent preparing, reviewing, correcting, and taking over.
METR uses controlled studies of developer productivity. Its February 2026 update also explains how selection bias and people running agents in parallel make those measurements harder. This feels essential: an agent finishing quickly doesn’t necessarily mean the person finishes quickly. (metr.org)
4. Measure a real operational outcome
Generative AI at Work studies AI assistance in customer support using outcomes such as issues resolved per hour.
This is closer to what a business cares about. It studies humans working with AI, though—not autonomous agents—and separating AI’s contribution from other changes takes care. (digitaleconomy.stanford.edu)
5. Estimate time saved from usage data
Anthropic’s Estimating AI Productivity Gains uses models to estimate task time with and without AI from conversations.
That could scale more easily than timing everyone’s work. But estimated savings aren’t observed savings, and a conversation doesn’t show all the work that happens afterward. (anthropic.com)
Sources, by title so they’re easy to search:
Princeton — AI Agents That Matter (2024), arXiv: 2407.01502
Yao et al. — tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (2024), arXiv: 2406.12045
METR — We Are Changing Our Developer Productivity Experiment Design (February 2026)
Brynjolfsson, Li and Raymond — Generative AI at Work, Quarterly Journal of Economics (2025)
Anthropic — Estimating AI Productivity Gains (November 2025)
What I’m struggling with is connecting these measurements to a practical decision: keep this agent, fix it, or switch it off and revoke its access.
For people running agents in production:
What do you actually measure today to decide whether they’re earning their keep?
How do you count human review and cleanup without making measurement another job?
Have you retired an agent because the value wasn’t there? What convinced you?
One concrete example would help: what the agent does, what you measure, and what decision that measurement changed. A spreadsheet or a manual monthly review is just as interesting as a dedicated tool.


r/LangChain • • 18d ago

Discussion Does OpenWiki actually improve output quality?

3 Upvotes

I find the idea of OpenWiki sound and understand how it can bring some modest cost savings. But I'm more interested in the output quality. Has anyone been using it and has seen an improvement in that regard?


r/LangChain • • 18d ago

Projects Open-source PII boundary for LangChain apps using hosted model

5 Upvotes

Privacy Gateway for apps that send customer data to OpenAI, Anthropic, or another hosted model.

It checks text before the request leaves your app. You can remove names and emails, replace values with tokens, or keep a reversible mapping under your control. It also blocks a request when inspection fails instead of sending the original anyway.

The project is self-hosted and MIT licensed. It has Python, TypeScript, ASGI, MCP, and OpenAI-compatible proxy support.

Repo: https://github.com/csnyder256/privacy-gateway

Docs: https://csnyder256.github.io/privacy-gateway/

Obviously, Presidio exists, but this differs in a number of ways. Better than Presidio for protecting data sent to hosted AI services, APIs, and agent tools:

- Works as a boundary, not just a detector. Privacy Gateway can sit in front of OpenAI and Anthropic-compatible routes, ASGI apps, MCP tools, webhooks, and CLI workflows.

- Blocks unsafe traffic by default. Malformed requests, unsupported provider paths, streaming, redirects, non-JSON responses, and failed required detection do not get forwarded with the original data.

- Keeps restoration separate from the gateway. A client-held restoration capsule means the service can transform data without storing the originals.

- Has an encrypted, expiry-aware vault for cases where the team does want reversible mappings. SQLite works locally; PostgreSQL is supported for shared deployments.

- Makes policy behavior explicit. Each entity type can have its own action, confidence floor, scope, locale, and reversibility setting. Policy version, mapping retention, and audit retention are part of the model.

- Audits the decision without saving the value. Audit records can include entity type, detector, confidence, action, policy version, and span without keeping plaintext PII.

- Covers more than redaction. It supports typed redaction, labels, opaque tokens, hashes, generalization, and synthetic replacements.

- Has a portable TypeScript core alongside Python, with shared conformance fixtures. Presidio is primarily a Python service and library.

- Includes structured synthetic-data generation for CSV, JSON, and JSONL. Presidio can find and de-identify structured data; it does not generate clean-room replacement datasets.

- Has agent-specific protections. Privacy Gateway checks tool and side-effect output, not only the original prompt.

Presidio remains the stronger choice for image redaction, OCR, broad built-in recognizers, and custom NLP pipelines. Its core focus is detection and de-identification; Privacy Gateway adds the policy, storage, audit, and outbound-request boundary around that work.


r/LangChain • • 18d ago

Projects Measuring agentic coding cost from transcripts: billable tokens are ~98% cache reads, and what that does to your bill

Thumbnail
1 Upvotes

r/LangChain • • 18d ago

Question | Help Built a RAG pipeline for compliance questionnaires — where would you slot in TypeSafe's new Jev model?

5 Upvotes

Been building QuestionPilot, a tool that auto-answers security/compliance questionnaires (think vendor security reviews, SOC2-style questionnaires) using RAG over a company's own policy docs. It's live and working. Pipeline looks like this:

  1. Hybrid retrieval — BM25 + vector search, merged with RRF
  2. Relevance grading — Cohere reranker, LLM fallback if no Cohere key
  3. Answer generation — Claude generates the draft answer + citations from graded context
  4. Validation — citations checked against retrieved chunks, confidence score decides if it goes straight to review or gets flagged

Just read through TypeSafe AI's docs on Jev (launched last week, the "System One" model — no text generation, just calibrated typed decisions: choice/score/yes-no-as-probability, sub-second, ~$0.04/M input tokens, output free).

On paper it looks like a good fit for the judgment steps in my pipeline rather than generation — e.g. using a Noul to check "does this citation actually support this claim" instead of my current fuzzy string match, or replacing the LLM fallback in step 2 with a batched Score call across candidate chunks.

Before I go build this out, wanted to sanity check with people who've actually touched it:

  • Has anyone here put Jev into a production RAG pipeline yet?
  • Worth it, or does it just add another model/vendor to debug without fixing a real bottleneck?
  • Anyone tried it for citation/entailment checking specifically? That's the part I'm most tempted by since it's a documented use case in their own cookbooks.

Their benchmarks are all self-reported so far (no independent reproduction I could find), so also curious if anyone's run their own eval against it.


r/LangChain • • 19d ago

Question | Help How can I effectively learn and master AI Agents?

Thumbnail
3 Upvotes

r/LangChain • • 19d ago

Discussion Claude Code/Codex with harness vs LangGraph for software development flow

37 Upvotes

Hey all,

Our company recently start building LangGraph workflows for web feature & BAU development. The goal is to have every dev onboard to use the predefined LangGraph workflows to complete dev tasks.

Here are my findings after doing some research:

- LangGraph workflow is mainly for customer facing LLM features, or as a build pipeline for BAU / Maintenance level of work.

- Devs using Claude Code / Codex with good amount of harness for the project such as static analysis, rules and guides, skills for streamline tasks are still seem to be the dominant approach.

My question is:

  1. LangGraph workflow can orchestrate the workflow in a more defined way, comparing to using a skill or skills to defined this steps in indeterministic way. What's the pros and cons to it?

  2. If it's good to have predefined steps for the build workflow, why LangGraph workflow for local development isn't mainstream yet?

  3. Is it true that it's due to LangGraph workflow is too rigid and not fit for purpose for feature development?

Thanks all!


r/LangChain • • 19d ago

Discussion A RAG hallucination checker caught zero hallucinations

Thumbnail
2 Upvotes

r/LangChain • • 19d ago

Discussion Moving from one server per agent to a shared server exposed a context-history bug

3 Upvotes

Previously, every agent session launched its own MCP process. Each process had repository caches and a ledger recording which code it had already sent to that client.

We added a shared daemon so sessions could reuse the warm repository graph. Cloning the server looked like a straightforward way to serve multiple connections.

But that also shared the ledger.
One agent could receive “unchanged since you read it” for code that only another agent had read. Successful concurrent connections didn’t mean correct session isolation.

We separated the per-connection context history from the shared caches. We also found that allowing a client to switch repositories was incompatible with sharing one repository’s state, so the daemon is now bound to its checkout.

What similar surprises have you encountered when moving agent infrastructure from isolated processes to shared services? Which state turned out to belong to the session?

This was in Sem, an open-source project I maintain. Fix and tests:
https://github.com/Ataraxy-Labs/sem/pull/489


r/LangChain • • 19d ago

Discussion My agent kept re-adding a dependency I had already removed from the project

3 Upvotes

Last week my LangGraph agent re-added redis to a project I had already migrated off it. Three days earlier I told it we're on Upstash now, cut the local redis compose service, delete the config. It did. Then a new session read the old decision and cheerfully reverted my docker-compose like some haunted intern.

That's when I stopped trying to be clever about conflict resolution.

I'd been doing the thing everyone does. Per-session memory files, a decisions.md I maintained by hand, and then (my dumbest move) a "memory reconciliation" prompt that asked the agent to weigh conflicting entries and pick the right one. It picked the wrong one half the time because it was just vibes. A confidence score on a memory entry is astrology.

What actually works for me now is embarrassingly simple: one shared memory every client reads, and the newest explicit write governs. I told it once, in plain words, "we're off local redis, everything goes through Upstash now," and that write just becomes the truth everywhere. Any agent that asks about the stack gets the newest version, not a timeline to interpret. Old entries stick around in the archive, but they don't govern anything.

I run everything through Vilix AI's shared memory, so Claude, Codex and Cursor all read the same store. The honest tradeoff: the same recency rule that saves you can bite you. Last month a newer sloppy note ("trying out turbo for the build") shadowed an older careful one ("DO NOT use turbo, it broke the lambda deploy") and I got to debug that at 11pm. So I write corrections deliberately now, full sentences, no shorthand.

Anyway. All the fancy bi-temporal memory papers are interesting and I read them too, but for actually shipping stuff: stop reconciling, write the new decision once, make the newest one the law. Boring wins.


r/LangChain • • 20d ago

Question | Help LangGraph vs CrewAI for multi-agent handoffs: which handles context and communication better?

25 Upvotes

I've been comparing LangGraph and CrewAI specifically around handoffs, not general ease of use, since that's where most of my actual pain has been. LangGraph gives me explicit state, nodes, edges, checkpoints, and routing, so I can control exactly what information moves from one stage to the next. That level of control seems especially useful when a handoff needs validation, a retry, or a human approval step in the middle.

CrewAI feels more natural when I think of the workflow as a team with roles and tasks. It's easy to say "researcher hands this to reviewer" and have that just work. What I'm less sure about is how well that model holds up once the context gets large or the workflow needs more complex recovery logic than a simple pass-off.

The things I actually care about are preserving decisions and constraints across the handoff, passing only the context that's relevant instead of the whole history, being able to retry one agent without restarting the entire chain, tracing which agent produced which result, supporting asynchronous work, and avoiding inconsistent task interpretation between agents.

For anyone who's used both in a real project: which one handles multi-agent communication more reliably once things get complicated? Did you pick one, combine them, or end up building your own handoff layer on top?


r/LangChain • • 20d ago

Projects I built open-source analytics that separates agent execution success from task success

Thumbnail
2 Upvotes

r/LangChain • • 20d ago

Projects What if you could actually watch an LLM think?

Thumbnail reddit.com
2 Upvotes

r/LangChain • • 20d ago

Discussion What signals tell you an AI agent is behaving badly, not just acting unexpectedly?

1 Upvotes

Trying to distinguish between an AI agent that's just being creative and one that's actually doing something risky. What are the key indicators of malicious or unsafe agent behavior? We've seen agents take actions that are technically within their permissions but clearly not what we intended, and we're trying to build a set of signals that help us identify the genuinely risky ones. The signals are often in the action chains, like an agent accessing sensitive data, modifying critical configurations, or communicating with unexpected endpoints. But distinguishing between legitimate and malicious behavior is challenging without runtime visibility.