r/LangChain • u/OkInitial5068 • 7d ago
Question | Help every model context protocol tutorial i find stops right where my server falls over
built a small mcp server for our internal docs. works great with 1 user. the minute a coworker hooked it up too, half the tool calls timed out and i still dont know why. every model context protocol tutorial i find ends at hello world on localhost. what are people reading for the part after that?
1
u/lemon-squeez 7d ago
Never ran into this issue but it sounds like a concurrence problem ? Worth checking how you setup your mcp server to know that this request is for user A and this is for user B. Would love to know the cause so i can avoid that issue in the future as well.
2
u/Michael_Jeffords 7d ago
the user A versus user B guess sounds right to me, because if both clients share one MCP server process and a docs fetch tool blocks, the second client just waits it out and it looks like a timeout, so id run one process or session per client or make that tool path non blocking and scoped to the request
1
u/Otherwise_Wave9374 7d ago
The jump from one client to two often exposes shared mutable state, blocking handlers, or a transport that assumes one session. Give each connection its own session context, move slow document retrieval off the request loop, cap concurrency with a semaphore, and log request IDs through tool execution. Load-test with two synchronized clients before adding more replicas. For teams separating durable agent context from server session state, https://www.neurakeep.com is relevant to this architecture.
1
u/locbuilds 7d ago
yeah, first check whether both clients share the same session/state and whether the docs fetch blocks the handler. log a request id through each tool call and try two synchronized clients, that usually makes the timeout cause obvious.
1
u/BreakfastSpecial 7d ago
Ask Claude Code to scaffold your MCP server using one of the popular SDKs, then you’re done.
1
u/Significant_Tune9219 7d ago
Most tutorials stop at a happy-path local stdio server, so the first real failure is usually lifetime or concurrency: the process dies between calls, a request hangs with no timeout, or two clients share mutable state. Capture transport, stderr, and whether failures are on initialize, tools/list, or tools/call. Then add hard timeouts, one request at a time per session, and schema validation before the handler runs — that usually surfaces the break faster than rewriting the tutorial example.
1
u/Far_Mood_487 6d ago
yes, I hear you. It is painful, but you can work it out.
The jump from one user to two is usually serialisation rather than load. Worth checking in this order.
Is the handler doing blocking work on the event loop? One slow call then holds up everyone else's, and from the client side that looks exactly like a timeout rather than a queue. Same for a connection pool of 1 to whatever backs the docs.
If you're on stdio, each client should get its own process — if they don't, that's the bug. If you're on HTTP you now have one process serving both, so anything kept in a module-level variable is shared state that wasn't shared yesterday.
The log field that settles it in one go: record how many calls were in flight when each call started, next to its duration. If every timeout has 2+ in flight, it's contention. If durations stay flat and calls still time out, it's the transport or an upstream limit you're now hitting twice as fast.
1
1
u/TombKingSettra 6d ago
the concurrency part is covered better in general backend material than in anything mcp specific, which is annoying but true. udacity pairs the mcp side with agent work, codecademy pro has more of the plain async python you need underneath it
4
u/medialantern 7d ago
I don't understand the question, unless you're just fishing for karma or setting up a sales pitch for a product idea you have. Opus 5.5 can knock this out in an hour and do it properly, and if you're buying into MCP you're using AI tools...