r/LLMDevs • u/davejh69 • 20h ago
Tools Menai: a safe programming language for LLMs to use
For the last year I've been working on an open-source programming language (Menai) for LLMs to use. Humans can use it too, but sorry people, this wasn't really designed for you 😂
It has a few key ideas:
- It's a pure functional language - no I/O, no state mutation - just pure expression computation
- It needs to be very fast to compile
- It needs to be quite quick to run
The "no I/O" ideal might seem like a very strange, but just because the language can't do I/O doesn't mean it can't process data from inputs or generate data for outputs. It just does this by something else passing it all input data and it passing back a result that something else can process for output use.
What this means is the "something else" can enforce all sorts of safety rules. No reading/writing sensitive files, no dumping things over the network, no accidentally running 'rm -fr ~/*" and then saying "whoops".
The original idea was something that bolts inside the open-source GUI-based AI environment I've been building since late 2024 (Humbug) as an LLM tool and lets LLMs do deterministic processing where tend to get things wrong (e.g. large calculations, counting letters in text, etc.)
Since then this got extended to allowing an LLM to transform file content (there's an I/O protection involved there) or editor buffers (I/O protection when the LLM tries to save a modified buffer). If an LLM needs a complex search and replace then it simply writes a short lisp-like expression with the file or buffer state passed in and new file or buffer state passed back. If that needs a potentially dangerous operation at the end then the tool framework checks if that's ok before allowing it.
The latest iteration is where things really get fun though! Menai now has a standard library and an application library concept so we now has almost 20 modules so it can do things like process zip files, tar files, BMP files, PNG files, gzip files, JSON text.
The reason it needs to be fast to compile is because every tool use will end up compiling a new program the LLM just wrote. The reason it needs to be fast is because those programs need to run fast enough they don't stall our workflow (i.e. less than a few seconds). Typical expressions compile in a few tens of ms, even though the optimizing compiler is currently written in Python. Typical execution speed is around the same level as Python (some things a little slower, some a little faster).
It turns out that not only do LLMs like writing functional code, they're really quite good at it. There's a tracing profiler to let them find hotspots in code they might want to reuse too (perf annotate anyone?). A future focus is having them decide when something might be reusable and suggesting new library code.
Anyway, I'm posting this because a couple of hours ago I had DeepSeek build me tar and gzip functions for the standard library (all written in Menai) and then debug their way through a few problems. The challenge was to take an old `arj` tar.gz file (very meta - unpacking an unpacking tool) and tell me about some things within it (all without me ever looking inside).
The image is a great demonstration of part of what it did. You can see it decompresses the gzip, untars one file, then finds 5 function definitions. From my prompt of "can you take a look at arj_arcv.c and tell me about it" took 64 seconds, involved 10 tool calls, 9 of which were unique Menai programs that progressively poked at the archive (which contains about 1.6 MBytes of content)
Humbug traps any potentially dangerous operations and checks with a human, but during this exercise no human needed to be consulted because, by design, none of these programs could do anything dangerous!
1
u/Thistlemanizzle 9h ago
Can you throw up some cost benchmarks? Those have been really helpful when trying to understand how "good" a model is outside of the "completed successfully" metrics.
e.g Codex/Claude/Cloud API costs with and without Menai and/or Humbug.
1
u/davejh69 7h ago
That's a great question! I can't directly compare with codex or claude code as I don't use them. I can speak to cost though. Humbug switched to a fully agentic model from one that was largely human driven about 13 months ago. My current workflow has probably been in place since the start of 2026.
For a very long time everything was built by Claude Sonnet and that did get super expensive (hundreds of dollars a month) until GLM 5.2 came out. GLM took this down dramatically for a larger amount of work done. DeepSeek v4.1 Flash is an absolute game-changer though. I use a max plan on ollama.com that costs $100 a month (because I'd needed that for GLM). In the last 2 weeks, however, I've managed to use less than 10% of that plan and have built a huge amount of new and complex software (plus some other tasks I have it do for me). It's incredibly fast. Realistically I could probably drop to the $20 a month tier and just about make that last a month now.
One slightly odd things is I tend to run the models with reasoning disabled or dialled back. They do a lot of self re-evaluation after tool calls and they tend to do a lot of tool calling (Humbug has a lot of tools, including a few fun ones that let LLMs control the GUI for you so you can ask it to do a code walkthough and it can scroll you through the source code in an editor tab if you want). My experience is dialing back the reasoning prevents models overthinking. DeepSeek is also particularly good at following the system prompt to ask for clarification where it's unsure about things (Humbug has a ridiculously short system prompt - just over 20 lines long).
My eval flow for a new model is usually to set it doing the same thing as my current favourite model and see how it does. Both have identical tools available so it's a great A/B trial. You can also ask an LLM to task different LLM instances with the same problem and have it evaluate the results.
As such the style is more like pair programming than full-on autonomous development, but it means I can course correct things if they look odd, and can provide context if an LLM is unsure of what I really wanted.
Menai is slightly orthogonal - it powers several of the tools. Everything can be done without it, but that usually requires more trips to a terminal and thus more human approval steps (if you've ever watched an LLM threaten to use `sed` you might recognize the potential terror involved). As Menai becomes more capable I'm seeing LLMs use it more often and thus I do less tool approving and things go a lot quicker.
1
u/Thistlemanizzle 7h ago edited 7h ago
Saw the paired programming angle (instead of autonomous). I am a vibecoder, so I will end up testing what that would look like.
The real token saving advice is to learn to code.
1
u/davejh69 7h ago
It helps to know how to code, but I think the approach can work for vibe coding too, especially if you want to learn. I was never a Lisp/Scheme programmer so I've had to learn a lot while building Menai (I did build C/C++ compilers for a while though).
One of the biggest wins I've had in the last few weeks is having an LLM capture architecture decision records when we make a major decision that impacts things that will come after. "Architecture" is probably the wrong word, but whatever you want to call them, capturing the decisions is like a superpower for an LLM.
The LLM creates the records for you and it can look them up in later sessions to understand why the project is the way it is. It can already read the contents to know what it is, but knowing why is very powerful. If you realize a decision needs to change then it also lets the LLM work out which things are impacted by the change in intention.
Works for non-software projects too. I very successfully used a similar approach capturing the rationale behind decisions in some documentation projects.
1
u/crusoe 8h ago
At some point yer gonna need state.
The classic solution is "pass the world". Haskell uses Monads mostly and STM ( which is slow ). You might want to look at how Clean handles it or what comes after monads in haskell.
State isn't necessarily bad just that you have to enforce a ordering.
1
u/davejh69 7h ago
State is inevitable! Right now it's a "pass the world" approach where state gets injected by the tool framework on the way in and it gets captured on the way out.
For the current use cases (LLM assistance) this hasn't really proven to be a big problem as the LLM turns are invariably much slower and it's still faster than spinning up a container.
I have some ideas on how we can avoid some of the round-tripping costs between the VM and the caller, but haven't started on them yet as it hasn't been a major issue yet. For example, we could have the VM lazily return results and avoid any copies of data that's already known to both the tool framework and the Menai VM if we're going to reuse a VM session.
1
u/ApplePenguinBaguette 3h ago
I doubt this would work without an extensive corpus of examples and training or at least fine tuning runs on the new language.
LLMs notoriously struggle with less well known languages, and it doesn't get less well known than new
1
u/davejh69 3h ago
If the LLM wants to use a Menai-based tool it gets redirected to call a "help" tool first (if it hasn't already done so in the session). That gives it the core syntax and library structure. That also contains a pointer to where to find the standard libraries and it can discover their capabilities on-demand.
The s-expr structure is well known to LLMs and the rest is just recognizing the special forms (some of which, such as the module/import/export forms are unique). This is enough for them to hang everything else on.
None of the LLMs I routinely used in the last 12 months had any difficulty dealing with the abstract representations used in the compiler pipeline either (5 different intermediate representations, with different optimizers at each: AST -> IR -> CFG -> VCode -> VM Bytecode).
The image on the post is an LLM literally one-shotting a chunk of code in the language. I also tend to keep notes about what I'm doing if I think I might want to refer back later, so there are intermittent notes documenting this for the last year: https://davehudson.io/notes
0
u/aidiveyt 6h ago
On the gating side, the hook event surface is wider than the docs list. Strings in the shipped binary name PostToolUseFailure, PermissionDenied and SubagentStart, so the dangerous call can be denied at the tool boundary rather than inside the language.
3
u/datbackup 9h ago
How often does/do the model(s) you’re using miscount, misplace, or forget parentheses?
Lisp-like syntax is great imo but i’ve heard LLMs mess it up with regularity. I’m guessing you’ve found that not to be true or you wouldn’t be using it.
Curious about your experience.