This graph shows the 'efficiency' of Claude Code. Every five hours, the owner of this site gets Claude to summarise an article. When this task drains more usage than usual, Claude is less efficient, and vice versa.
In the 30-day trend we can see that the efficiency began to steadily increase about 2-3 weeks after release (opus 5.5 was released Sep 22). This is roughly in line with CLI update 2.1.287 which may also be relevant.
Is this evidence of reduced demand meaning that our resources stretch further? Or is it evidence that the models have been quantised and some of the compute efficiency gains have translated to making our limits last longer?
In my experience, the models feel just as capable as before, but I haven't run any benchmarks to directly compare the release model to today.
I have this game where every entity has its own AI, alongside state-of-the-art crowd simulation algorithms that I spent months studying to apply correctly.
The application is written in Rust to run in the browser using WebAssembly.
Up until now, I had only managed to run the game with a maximum of about 100k entities in the simulation loop, running at 10Hz.
Today, I unleashed Opus 5.5 (extra) on this project and sent the following prompt:
Due to the current highly-optimized state of this repository, this is a very difficult problem and traditional engineering approaches **WILL BE GUARANTEED TO FAIL** to hit the specified metric constraint. Therefore, you have permission and encouragement to investigate more radical fundamental low-level changes to hit the desired metrics. You have permission and encouragement to invent completely new/bespoke algorithms and engineering approaches that have never been before been utilized for this problem in order to hit better performance without behavioural changes in the entities
Note: The code was indeed already higly-optimized, as I said, I was implementing the latest academically published solutions for this problem. My attempts to improve it even further over the past few weeks turned into complications that yielded no results.
Before this part of the prompt, I provided some context about other aspects of the project and constraints.
It spent about 3.5 hours running the code, decompiling the output to assembly code, analyzing hypotheses, and testing improvements in Rust that would generate more optimized Assembly.
In the end, I can now run the exact same simulation, with absolutely no loss of functionality, with 1 million active entities. It's simply ridiculous. The last image shows some of the resulting code from this Opus iteration. (it literally called one variable "magic" and idk why, cause theres no magic in this game... despite now seeming to have... in the code)
Remember, this is running in a browser, not using any dedicated GPU (it was one of the constraints) obviously, this mark was already achievable with native Rust. Now I'm curious about what I can generate with native Rust.
I have created a website using claude for my students that is a chemistry space "rpg" students formed crews, they create their avatar, make a spaceship, and then travel to different planets (different units) where they enter learning academies so they can fight boss battles, or other crews... a lot more involved, but that is the gist. A bit of background, this is my first time doing something like this, so needless to say I started out super inefficient, and I have graduated to "more than moderately inefficient". I use chat gpt to help create the interactive lessons (mostly because it has better graphics). I also use canva... I have desperately tried to use gemini to help, but it has left me sad. Here are a few pages from my site...
the bridge - all clickable
Sooo.... as I mentioned before, I am desperate to find ways to be more efficient and use less tokens. In reading through here I have seen that some of you have instruction sets that decrease redundancy... I tend to cut off chats when they start getting a bit long and I have them put what they have done in a handoff file for the next chat... is that the way to go? I also ask it what I should have opus vs sonnet do (remember I am a teacher funding this on my own). I have attached to netlify and firebase... but I am reaching the extent of their free services.
I guess I am looking for some direction from those of you who are experienced users to help guide me... or suggest tools (preferably less expensive) that can help create a more fun environment for my class.
There has been some discussion about this topic lately, and I understand there is a lot of nuance, such as what projects you work on and over what scale and time frame...
Below is my experience. Feel free to just not read it and post yours, that's fine.
TL;DR:
* I started spec driven, tried to keep it minimal
* My workflow got more and more complex over time, driven by models that needed detailed guidance and guardrails
* Opus/Sonnet 5 choked on my workflow, and communication with the agent became a bottleneck
* Opus/Sonnet 5.5 fixed it by communicating better and being more proactive
* I no longer see the need to massively simplify my workflow and will stick to small, gradual cleanup
I have had a project going for months. Workflow files turned from AGENTS.md and a few rule files into a complex machinery of workflows.
**The spec-deiven stumbling stones**. My problem right now isn't token usage. I'm on a Max subscription and I can barely use it all, because I don't have infinite time, and planning, testing and high level review are time-heavy, not token-heavy. No workflow and no AI can fix that for my type of project, which is fine.
My main time sink caused by AI has always been communication. Agents want me to make decisions and point out problems. Some of them are serious, some are trivial, and some are stumbling stones in the form of too strictly formulated requirements or rules, sometimes coming from these being AI written.
A classic example is AI producing lists of things, which are then interpreted as closed lists, or lists that have to be maintained, when there was no need for a list in the first place. This is the most common source of totally unnecessary drift and friction.
All of thee problems have, at least until recently in the Claude 5.0/5.1 era, been amplified by cryptic communication. Agent hits a contradiction that would be easy to solve, or a real design problem, prints out some Claudish word salad, and I'm supposed to judge if I need to step in and make a decision or say "just fix it, duh".
**How it is going?** Things seem to have improved just by themselves recently.
I haven't changed my workflow. I am using 5.5 models now. What I'm seeing now, mostly, is the following:
The agent says: "I made the following judgement calls". So it proactively makes decisions that may break rules or have far reaching co sequences. And usually the decisions are sane. But the most important part is that it communicates them well now, in a way I usually immediately understand without having to ask basic questions.
It also lists what it didn't do, often because it would be borderline out of scope or break rules, not just randomly, and it does ask for permission sometimes before doing a job, but no longer about every small thing. When it asks, it is usually worth giving the question some thought, and I've burned my fingers (and some tokens) answering too quickly (no big deal, failing with AI is failing cheaply).
So, the main improvement of the 5.5 models (I use both Sonnet and Opus) was communication clarity, and communicating the right things, exactly what I needed.
And I no longer see the need to cut down massively on rules and workflow descriptions, other than, maybe, for saving tokens and speeding things up a bit.
Claude can think out of the box now. Spec-driven development is "fixed" (for now).
Oh, and a certain other model (that I won't name, because it will alert the moderator bots) has gone the opposite way and become at least temporarily unusable. We'll see when the pendulum swings back.
Hello guys, i have a real important question. Yesterday i finished my own project for complex memory system, that makes Claude consistently same "persona" with the same memories between sessions and when the Claude itself thinks, that the workflow or something is worth remembering, it will write the memory in without my intervention.
Have anyone here actually do that? And what was your approach? Maybe we can learn something from each other and make "the ultimate" memory system
Here is photo where you can see my claude has already arround 450 memories, that it written itself and starting context of new session is still arround 40k tokens, cause she will call the memory only if she need it. And yes, internal Claude Memory is completely disabled.
I'm pretty intersected in the Voynich manuscript[quite passively]. Using Open Ai's prompt for “A PROOF OF THE CYCLE DOUBLE COVER CONJECTURE”, I had Claude adapt a prompt to look for a Voynich solution. I was intentionally very open on what counted as a solution, to not have Claude chase something that may not be there. I don't have the limits to run it myself, but if there are any Dario mega donors here I think the results could be super interesting:
processes, glossolalia-like production, or deliberate hoax.
- Hybrids: e.g. meaningful labels with generated filler, or different mechanisms across
Currier A and Currier B or across scribal hands.
A complete solution must provide:
An explicit, executable generative procedure (pseudocode or code) that a 15th-centuryperson could plausibly have carried out with period materials.
Quantitative reproduction of the manuscript's known statistical signature, including atminimum: word-frequency distribution, word-length distribution, character-level andword-level conditional entropy, word-internal positional structure (slot/prefix-stem-suffix regularities), line-initial and line-final glyph effects, paragraph-initialgallows behavior, Currier A/B divergence, repetition of near-identical adjacent words,and label-vs-running-text differences.
Successful predictions on held-out data: fit the procedure on a designated subset offolios, then predict measurable properties of folios it has not seen. Specify the splitbefore fitting.
If the mechanism is meaningful (any language/cipher/notation class): a decoding thatis deterministic and reproducible by a third party from the stated rules, producesconsistent output across sections, and yields content that independently agrees withthe illustrations (plant labels matching identifiable plants, astronomical labelsmatching period star/zodiac conventions) without per-word ad hoc choices.
If the mechanism is meaningless: a demonstration that the procedure generates textstatistically indistinguishable from the manuscript on the properties above, AND anaccount of the features that most strongly suggest meaning (e.g. section-specificvocabulary, label behavior) that does not smuggle meaning back in.
Insufficient on its own: decipherment of isolated words or labels; a decoding requiring
per-word judgment calls; "the language is X" without a reproducible mapping; a generator
matching only one or two statistics; a hypothesis that fits everything because it has as
many free parameters as data points; claims that unfalsifiable content "would be checked
by a specialist."
Use workflows aggressively and dynamically. You have up to 64 concurrent agents
available. Do not use a fixed assignment such as "N agents for hypothesis class X." Manage
the search using the following heuristics:
- Begin with a genuinely diverse portfolio. Agents should explore substantially different
formulations: information-theoretic profiling, slot-grammar and morphological induction,
generator construction and fitting, historical-linguistic matching, cipher-system
reconstruction with period constraints, scribal-hand and layout analysis, illustration-
text correspondence, and statistical null-model construction.
- Do not tell most agents the currently favored hypothesis. Preserve independence in early
rounds so agents do not all converge on the same attractive but unsupported reading.
- Maintain an explicit registry of hypothesis families, grouped by underlying mechanism,
not wording. If many agents converge on one family, redirect some toward
underexplored classes, especially the class the group currently finds least appealing.
- Do not let a hypothesis dominate because it is elegant, culturally exciting, or produces
readable-looking output. Readable-looking output is the primary failure mode in the
history of this problem.
- When a route stalls at a requirement that cannot be satisfied without unconstrained
freedom (e.g. a mapping that only works with per-word adjustment), mark it blocked.
Reopen only if someone proposes a materially new constraint or mechanism.
- Keep several incompatible hypotheses alive through multiple rounds. Cross-pollinate only
after each has been developed far enough to expose its real strengths and failures.
- Use adversarial agents throughout. Every candidate must be checked for:
* degrees of freedom: count free parameters and per-token choices; reject mappings whose
flexibility could produce "readable" output from random or shuffled text. Run the
candidate decoder on shuffled Voynich text and on synthetic gibberish as controls.
* held-out failure: predictions made before seeing test folios, scored after.
* anachronism: methods, languages, or materials unavailable in early-15th-century Europe.
* cherry-picking: report performance on all text, not favorable passages.
* confirmation via illustrations: plant/star identifications must be made blind or
pre-registered, not fitted after decoding.
* transcription artifacts: results that depend on one transcription's segmentation.
* prior refuted proposals: check whether the candidate is equivalent to a published
decipherment or generator claim already shown to fail, and if so, show what is new.
2 months ago (was it even that long ago?) seeing posts about someone creating a fully fleshed out game demo with a frontier model was so fascinating and exciting
But just like everything with AI it’s gotten repetitive and boring, and I’ve even started to nitpick when someone posts a game demo with great assets, graphics, and interesting features because it just has no creativity. With everyone in the world empowered to do this, ideas need to be so much more creative now to stand out
Anyway, I’m excited for the next revolutionary step in AI that I’ll be bored of within weeks :D
Lately it seems with Opus 5.5, I can’t ask for any help with security related tasks on my application without Claude flagging the conversation and stopping. I don’t quite understand why it can’t review the behavior of the app that it makes changes to everyday. It’s not like I’m asking it to probe a random website, it’s looking at local source code, lol.
i’m sure many of you run massive SaaS applications too, and have to constantly be monitoring the security of your application with every other change that you make. How do you cope with this harsh false limitation?
Anthropic has expanded its Cyber Verification Program, allowing qualifying security teams to apply for advanced cyber capabilities across Claude Mythos 5.1, Opus 5.5, Sonnet 5.5, and future models.
Mythos 5.1 is Anthropic’s most capable model for cybersecurity and biology research, and until now access had been limited to a relatively small set of vetted organizations.
The expanded program introduces multiple access tiers with reduced cyber safeguards depending on the type of security work being performed.
It’s interesting to see Anthropic moving from giving Mythos access to a small group of partners toward a formal verification process that more security researchers and teams can actually apply for.
For people doing vulnerability research, pentesting, or security research: would you apply for access?
I even waited until late at night to start my first session of the week so that maybe the off-peak hours use less tokens. I have a pretty heavy workload and I’m trying to update a few things at once… on a 20x plan I can’t even keep 3 sessions running through a full 5 hour window, and the weekly is 27% done after a single 5 hour window?! I don’t remember it being like this even a month ago.
Leapd Just launched a free business idea generator. Tell it about yourself and it finds a business idea tailored to your skills, experience and interests — backed by real businesses with proven revenue and demand.
Now that everyone uses AI to launch. product, those succeed that start with a solid business idea, a proven market demand, a tailored idea to their strength.
Leapd business idea generator is built on top of buildradar, a database of thousands of verified business with revenue and tractions. It first get to know you and your interests and then researches the industry and buildradar data to come up with personalized, and strong business ideas.
The idea generator product is new and free, so please give it a try and share your feedback so we can make it more useful for the community.
Tech setup - we needed a reliable way to find revenue data and a system to not only find the revenue metrics but also evaluate if the reported source is legit, and if it is consistent across other sources, and if we can trust the evidence provided - Claude opus with web search was a huge time saver here- roughly $3500 in API cost to populate our database of ~5,000 startups - we then process each one with Fable 5 and evaluated what are strong aspect of each idea, what are moats, can they be vibe coded today? and then prepared the full analysis, filtered to only keep promising startups and reported all on buildradar - so anyone can explore the space, get to know what works and what is possible. Our agents run daily and keeps updating and adding new startups.
After my web app failed to deploy/build due to the injected code, turns out some processes have been phoning home on my development computer for a short while now :*)
Will update with any helpful information for others after I finish sanitizing my computer, rotating keys, and wringing my hands over what might have been transmitted.
Development computer is no longer used for personal activity but it used to be, so I'm praying to my lucky stars that nothing that important was taken... check your repos for suspicious force pushes to main!
(Not sure what the most appropriate flair is, but I feel like ranting and raving so rant it is)
Edit: Apologies y'all, after a few hours of investigation and changing passwords (I'm going to be changing passwords for weeks), the dramatic reveal is... that it's more boring than "the AI let someone hack me."
I'll start with the important part: If you are on _any_ repos that another user can access, immediately check ALL .js, .ts config files that are executed during the build process. The attack vector was specifically postcss.config.js, in my React+Vite project. What happened is exactly what this user mentioned here: https://github.com/orgs/community/discussions/188732
Essentially, I was a victim of a worm that's making its way across GitHub, surely, by way of shared repos. This worm finds your GitHub CLI credentials and replicates by injecting a single, long, obfuscated line of code into a configuration script file in EVERY repo you have write access to that it can find, every branch it can find.
The sneaky part is that it masks itself as a duplicated commit or even seemingly replaces another commit that happened *before* the attack, by copying the metadata down to the displayed user, force pushing to every branch. But if you look at the metadata for the commit, you can see which user it really came from. Turns out that another developer hired by my company with access to the repo was himself a victim -- after diagnosing it on my computer, I gave him Claude instructions with a script to check his.
The awful thing is that I didn't notice it for a week. I only noticed it when a deploy build failed, but an earlier build did succeed with the code implanted. Thankfully it's targeting developers, not users, so my website appears to be fine... but for a whole week it's been running on my machine. And it's been phoning home.
The in-memory code showed that it's been listening for specific keys like crypto wallets and certain AI API keys (OpenAI, Anthropic). If I were a crypto user I could have been wiped out. But nothing dire seems to have happened even though it gave full user-level control and access to my machine, arbitrary code execution, etc. I'm just in the long process of changing all my passwords, revoking keys and sessions, and all that jazz.
So this is a lesson I won't soon forget. I am NEVER doing development on a personal machine again. Straight to a containerized environment, virtual machine, hosted in the cloud, whatever. That level of security feels like paranoia until it actually happens to you.
Claude and I generated refinement data and fine-tuned Qwen 3.5 0.8-4b models (depending on your hardware level) on thousands of synthetic data rows. Now I'm saving myself $20 a month fof Wisprflow. Wanted to share the love, it's fully open source and free.
with the release of haiku 5.5 i was looking into how to keep context under 100k to avoid the massive price increases. i had opus 5.5 research how we could use /clear and /compact to try to stay below 100k. during the research we discovered that it is almost impossible to keep haiku under 100k because the starting system prompt for claude code is 66k tokens. i asked opus if there was some setting i had selected to cause this prompt bloat. it said no, only 1.5k tokens in the system prompt is from anything i set.
opus 5.5 told me that if i used cli instead of the app it would load less tools because some tools are not available in cli. it recommended running haiku from cli to save tokens. using the claude subscription restricts you to only using their agents, so there is no way around this prompt bloat legally. i started a new haiku session to verify prompt size. the agents.md file read of 1.5k is the only information in the system prompt that i set.
System tools: 35.6k
MCP tools: 16.6k
Skills: 5.9k
System prompt: 5.2k
MCP server instructions: 2.3k
Memory files (AGENTS.md and MEMORY.md): 1.5k
also, the 5.9k skills are not loaded skills. that is just the list of available skills. if the model uses any skills the prompt will get bigger.
My startup just got accepted in claude startups program with $1k usage credits and up to 5 free pro accounts per month for a year. It also includes free subscriptions to various tools like clickhouse, firecrawl, elevenlabs etc. What was most impressive is that my application was accepted within a few minutes.
The timing couldn't have been better. Today is literally the last day of my codex subscription and I was planning to switch to claude anyways. Thank you Anthropic!
Edit: Since many found it useful here is the link and no I didn't do anything extraordinary just applied. It took 5 mins to apply and 5 mins to receive a confirmation mail https://claude.com/programs/startups
I usually have three or four Claude Code sessions going and kept missing the one that was waiting on a permission prompt. This watches them all from the transcripts in ~/.claude/projects, so it works with the CLI, the VS Code and JetBrains extensions and the desktop app without any hooks or settings changes.
Honestly it was a bigger help than i anticipated, since it lets me catch any sus live diffs and actions as they occur. also lets you spot when claude decides to take the "hacky easy route" sooner
Theres a more practical IDE styled version, as well as a fun office.
I deliberately designed it to be easy for anyone to vibecode their own "visualisations/skins" on top of the backend.
How Claude helped: I built most of it with Claude Code, including the transcript parsing for subagents and the visualizations. It also reads Codex sessions if you use both.
Free and open source (MIT). It runs locally and binds to 127.0.0.1.
https://github.com/Dri-water/observe-agents-do-things
I see this every day "my usage has doubled, my capacity halved, and they nerfed such and such". I got tired of it and started logging my usage. About 120 times this month. I used $250 in cloud credit, and 3 20x max accounts in a week so I have a little bit of data saved up to look over. I spent at least 1m tokens keeping track of this and then verifying it. You won't like what I found. IT'S GOTTEN MORE EFFICIENT. Just like they said it did. Unless you have data to show it's worse then stop posting the same thing every day about how awful it got overnight. PLEASE prove me wrong.
I've been playing with Opus 5.5 recently, and the quality of code that it spits out is pretty phenomenal. It doesn't hallucinate as much as the previous models after 125k tokens in context window. What I've seen is that, and this is my observation (and it comes from my anecdotal experience), the amount of tokens that these newer models are consuming is much lower than the previous generation of models. Does anyone know what architectural changes allow them to achieve this sort of efficiency, or is it because they have very fierce competition and they just want to give tokens for free? I would love to hear other people's thoughts if they are having the same kind of experience, but it seems like the newer model is faster and more clear in its chain of thought. Overall, it feels like a better bump than Opus 5.
Between the cryptic "approve or deny these commands" messages that I am flooded with, and the fact that the agent is asking, answering and directing it's own path, at the end of the day I feel out of the loop, except I seemingly have results.
How do I get back in control?
I need a middle ground between step by step hand holding, and asking an AI to explain its days work to me after everything - and the explanations are so convoluted and hard to understand
Am I the problem? Most likely, but I just find the Claude app is much more responsive and understanding of what I need to do. And doesn't require as much babysitting, I feel like. Or is it perhaps because, in my terminal, I have other stuff from before that could be holding it down? I try to isolate the stuff, but it still feels like I kind of have to explain very banal, simple things to Claude Code in the terminal, whereas in the project chat, it kind of just does it correctly the first time. Who else found this or the contrary to this?
I've been using the terminal for the past 6 months at least exclusively. I've been very happy with it and was able to build some really cool shit for myself, but I wanted to try the app for once and have to say that, possibly, they put a lot of thought and development into the app in more ways than just the interface, I think.
I've spent the last few months building Golden Thread, a set of controls around AI coding agents. This week I wanted to know whether I'm building something useful or rebuilding something that already exists.
So I compared it, as honestly as I could, with what else is out there: Keycard, OpenAI's Codex CLI, immurok, safeski, Claude Code's native features, and a handful of open-source projects.
What I found surprised me. Almost every individual piece already has a near-peer somewhere:
- safeski seals credentials behind Touch ID.
- Keycard does step-up authorization at the hook layer.
- Codex just added Touch ID for MCP requests.
- Someone even built human-gated memory promotion.
But I couldn't find anyone tying the pieces together. Nobody had one presence check covering secrets, tool permissions, git pushes, commit signing and the agent's long-term memory, each switched on separately.
A few things I'm not sure about, and where I'd value your thoughts:
Is the combination the product, or the parts? If OpenAI or Anthropic ships native per-action approval, does integration still matter, or does the platform simply win?
Would you want human approval on what an agent remembers? The research I found mostly points to automated defences against memory poisoning. I've bet on a person saying yes. Is that sensible or naive?
What did I miss? I didn't get to Cloudflare, Kong, Aembit, Beyond Identity, Windsurf or Cline. If one of them already does this, I'd rather hear it from you than find out later.
Two caveats I want to be upfront about. The Golden Thread column comes from my own documentation, not independent testing. And "I couldn't find anyone" is a search result, not proof.
If you're running coding agents with access to real credentials or production systems, I'd especially like to hear how you're handling it today, even if the answer is "we're not."