Rendered at 16:44:28 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
pimeys 7 hours ago [-]
The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between the models, and Kimi K3 was the first one that actually could compete or be even better than Opus or Sol in our use cases, with a fraction of the price. All our developers use K3 as their programming model, and it now powers a big part of our systems instead of Opus and GPT. Surprisingly the new Sol pricing is quite similar to K3...
Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.
miki123211 5 hours ago [-]
It's funny how quickly we went from "the greedy US companies are subsidizing prices to keep competitors out of the market" to "the greedy US companies are overcharging because they are greedy."
fluoridation 5 hours ago [-]
They are subsidizing the non-API use cases and overcharging on the API use cases.
pimeys 4 hours ago [-]
It's interesting if they need to cut off their subscriptions to be able to compete in API prices. Very interesting...
edg5000 3 hours ago [-]
I'm not happy about it, but API prices seem actually reasonable if you compare it to free-market pure inference providers. They need to buy the same power, same hardware as the big labs without needing any R&D. This means that the coding plans must be subsidized. I'm not happy about that, because it means we'll be paying more, until hardware and maybe power drops in price, which, if we look at e.g. the housing market, may never happen during my lifetime.
senordevnyc 5 hours ago [-]
No, almost everyone has been in agreement that they’re subsidizing subscriptions, but there have literally been dozens (hundreds?) of threads on HN in the past 12 months with people vehemently arguing that API prices are subsidized.
jmuguy 3 hours ago [-]
I think both can be true - they're losing money on the API and they're still charging too much. Which is really where the music stops for American AI investment. I also use K3 now and its perfectly capable for the development work I'm doing. I don't shed any tears for OpenAI or Anthropic.
platinumrad 2 hours ago [-]
API are not subsidized. We know this is true from the existence of independent inference providers (a lot of which are crypto companies that would otherwise just be mining if inference weren't actually profitable).
jurgenburgen 2 hours ago [-]
Those providers are VC-funded, they don’t have the leeway to pivot away from AI. I would be interested in some examples though.
Barbing 1 hours ago [-]
Any public company providing inference who’s claimed its profitability in a way they risk legal penalties if untrue?
(Makes sense to me, just curious about this additional piece.)
lucassz 29 minutes ago [-]
Certainly not everyone. I find it far more likely that both subscriptions and API pricing are profitable in their own right. The fact that some people get great value out of their subscription (just like some people get great value out of their car insurance) doesn't mean it's subsidized.
senordevnyc 7 minutes ago [-]
That’s a fair distinction. I think they are willing to subsidize individual subscriptions for people who maximize usage, but their overall subscription business may be profitable.
Der_Einzige 14 minutes ago [-]
And every one of them is wrong.
5 hours ago [-]
conjecTech 1 hours ago [-]
Both can be true simultaneously. American companies were comfortable with a high cost structure because it padded their revenue numbers. The government tried to block out competitors through import/export controls, and it ended up backfiring by creating more efficient competitors. Very similar to our competition with Japan in the 80s. We'll see if we give the AI labs the kinds of sweetheart protectionism that US car makers got.
phoghed 1 hours ago [-]
Can you explain better how OpenAI, for example, would simultaneously subsidize tokens while overcharging for them? Not sure I get it.
solarmist 33 minutes ago [-]
As one example: Different access patterns. If you're using the app and have a subscription you get cheap tokens in the hopes you need more at which point you pay API prices which are much more expenseive.
phoghed 17 minutes ago [-]
I agree with that. However, there’s much dialogue around subsidized tokens for business use too, that people paying for the tokens are also vastly underpaying vs the “real” cost. I certainly don’t know the answer to that. Maybe the Chinese companies are also doing it. Maybe nobody is doing it.
Looking at reserved capacity cost for PTUs on azure, which I think they’d probably not subsidize but can’t be sure, I’m inclined to not agree with the vast undercharging for tokens hypothesis.
square_usual 23 minutes ago [-]
It's simple, they aren't subsidizing tokens at all.
Der_Einzige 14 minutes ago [-]
True for api pricing across the industry. Not as true for 200$ a month plans assuming you used every bit of it.
pfisch 19 minutes ago [-]
I don't think that kind of protectionism will work with a digital asset. It is much more difficult to erect barriers when there are no physical goods and transportation across geographic borders is instant and free.
conjecTech 8 minutes ago [-]
It will likely be on the enterprise side if it happens. Hard to control supply, but if demand is concentrated to a few hundred entities, that's easy to do.
gagan2020 6 hours ago [-]
I shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.
Bayart 5 hours ago [-]
Try DS through their own API if that's feasible, AFAIK they're much cheaper than through OS due to cache hit rates.
bogdan 4 hours ago [-]
Is there actually a difference in cache rates between OR and official API? I have a preset set up on OR so that I only send traffic to deepseek. The preset is important otherwise you will send traffic to different providers but if you weren't doing this already then what can I say, water is wet, of course cache rates will be awful. I get about 70% cache hit rate with the preset which is appropriate for what I'm doing. I haven't used the official API though.
irthomasthomas 3 hours ago [-]
Zenmux say the cache hit rate is 98% for the deepseek flash API. I don't know why, but performance is definitely worse using openrouter.
Openrouter is going to cost you a lot more than the 5% fee, unless you lock the provider.
tkgally 6 hours ago [-]
I was struck by a video ad that Google released yesterday with testimonials by three developers about using Gemini 3.7 Flash [1]. The point they emphasize most is price, followed by latency. The marketing strategy definitely seems to be shifting.
You play to your outs. Gemini is far behind on quality.
eru 6 hours ago [-]
> And these models are not going away, nor their prices going up [...]
Well, DeepSeek just raised prices.
pimeys 6 hours ago [-]
And Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.
pimeys 47 minutes ago [-]
They did now. Landing somewhere between Terra and Luna now per task, with the quality of Gemini 3.7 flash.
1 hours ago [-]
ywvcbk 7 hours ago [-]
> Opus or Sol in our use cases, with a fraction of the price.
I assume it's highly use case dependent, though?
Even before the price cut seems like Sol was price competitive with Kimi
Long-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than Opus right now.
You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.
mrngld 6 hours ago [-]
China's 50 Cent Party being a real and noticeable thing (and the two biggest things they like to shill is open weight Chinese models and the futility of resisting a Taiwan invasion), I have to take things like this with a healthy dose of skepticism without corroborating data, since independent evals didn't show the price per task lead you're showing.
If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.
pimeys 5 hours ago [-]
Internal reports from company? Maybe not. I'm just saying you have to eval eval eval if you are working in this industry. There's a ton of victories in price, and price is right now the key thing all the customers are talking about.
It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents.
If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive.
And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.
c16 4 hours ago [-]
> I'd say a healthy 75% cut to token prices and they are back in competition.
Surely you don't want them to be the reason the bubble bursts?
pimeys 4 hours ago [-]
Yep. It's a bit scary also. There's a lot of opportunity in the market now, but the downfall of the big US inference labs is going to hurt here in EU too sadly...
gloryjulio 32 minutes ago [-]
We know the price wars are coming. It's gonna be interesting to watch how it plays out in real time. Anthropic/Openai probably want to race to the IPO before the price wars start to have real impact.
hacompanian 1 hours ago [-]
Not including Grok/Cursor seems a major flaw in your model
pimeys 52 minutes ago [-]
Oh we did try to get Grok for evals but they had some weird EU limitations last time we checked. Which the open weight models don't have.
ee334y5rthsrth 5 hours ago [-]
Tell me why price adv not working this same on web pages. Why price od advertisment on portals, social media etc. not fall down?
netsec_burn 16 hours ago [-]
After using Claude for a long time, I tested Sol 5.6 for the first time today. Love it, its an incredibly capable model and uses far fewer tokens/time thinking. Its what I imagine Fable would be if I haven't been downgraded on every conversation - even after completing the verification program. I think I may cancel my Claude subscription finally.
jchw 15 hours ago [-]
I think Fable's dominance is overstated. It definitely has the lead, but quantifying what that lead actually is is really hard. I'm using GPT 5.6 Sol to do some shit that I personally would consider "crazy" - low level undocumented hardware driver alchemy, reverse engineering highly obfuscated code, even a bit of screwing around with a rendering engine in Vulkan, really just about the most complex tasks I can get any model to do, and it does great. For the more advanced stuff, it definitely needs the effort bumped. But even with the effort bumped, the token usage really doesn't seem to skyrocket too badly until at least you hit xhigh and max, which really only seem to be necessary if you are doing genuine crazy stuff, so it's not that bad. I did similar stuff with Fable. In fact, I went directly from an Anthropic subscription with Fable to an OpenAI subscription with Sol, more or less, and it really felt pretty seamless. If anything, I was thrilled to realize how much I actually preferred Codex CLI, to the point where I started using it at work too.
Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).
However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.
andrewingram 14 hours ago [-]
I used Sol to extract the remaining decryption keys from the Super Mario Maker 2 (Switch) game files. Someone had previously extracted all the keys from the original release, but not any of the new ones from updates. Not only did it succeed, but it helped me understand the data sufficiently to add support for “Super World” rendering to my level viewer (which I made back in 2021), eg the little widget at the top of https://www.smm2-viewer.com/players/B16-306-GVG
I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.
mrngld 5 hours ago [-]
Interesting. I've been wishing that old 'Stars!' game from the 90s would play easily on modern systems. I'd love it if we'd got to the point where I could point Codex at a folder with an ISO from my CD of the game and tell it to go reverse engineer it all for understanding of game mechanics, then go recreate it in a modern language capable of running cross-platform. Scarcely any need to improve on graphics, it could even be a PWA.
There's people that have tried to contact Jeff McBride and follow the IP trail but the IP is currently owned by a company that went defunct. Not sold, but no one is even bothering to register its LLC any more, it's simply dead.
mwigdahl 3 hours ago [-]
Stars! was awesome, many fond memories of play-by-email games of that...
mjhagen 10 hours ago [-]
I was having it look at creating a driver for some old scanner and it actively looked up exactly where that gray area for my country was wrt decompilation.
alchemist1e9 13 hours ago [-]
Fable is almost unusable for anything but super boring mainstream stuff. I was getting safeguard flagged so often I’ve significantly reduced my usage out of fear they will blacklist/ban me.
Some of the topics it’s flagged have been hard for me to understand what it seeing that can be remotely concerning in my requests.
krull10 7 hours ago [-]
I cancelled my Claude max subscription. Somehow every query I sent was flagged as bio or chem, even pure mathematics questions. Not going to waste money paying for a “max” subscription that won’t ever let me use the top tier model…
Sol is great and has never blocked a request, and generally gives great answers. Happily switched over to it now.
andai 12 hours ago [-]
The safety is really funny to me. I ask it a lot of extreme stuff and it goes through, but I ask it mundane stuff and hit the filters all the time.
vjvjvjvjghv 10 hours ago [-]
Reminds me of the URL blocking of my company. Nature.com is being blocked but I can access a ton of super sketchy download sites.
vintermann 11 hours ago [-]
It has learned a little too well how it works in human society.
xscott 12 hours ago [-]
I've gotten flagged for asking questions about tokens and tensors. That makes me believe it's not about safety, it's about protecting their turf. I cancelled my subscription - same fear about getting flagged too much leading to a ban.
vineyardmike 7 hours ago [-]
They said they also block usage of Claude models to build ML models.
Which is definitely protecting their turf, but also probably a little bit hiding their “RSI” abilities for competitive reasons. My theory is that a lot of “safety blocking” is actually WIP training of new business directions. Anthropic has started hiring biologists and has opened a preview of a “Claude code for bioinformatics”. I’m guessing they’re tweaking their bioinformatics market play, and block “bio safety” requests so competitors can’t learn about their training.
xscott 13 minutes ago [-]
What's the distinction between "protecting their turf" and "competitive reasons"? I see them as the same, but I could be missing something.
That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.
KronisLV 8 hours ago [-]
OpenAI and Kimi are both pretty okay alternatives! I guess GLM 5.3 on Max reasoning as well but for more limited domains.
alchemist1e9 4 hours ago [-]
I agree. It’s flagged me on discussing fast GEMM implementations for large regressions, a discussion on theoretical physics math, a discussion on designing a type of
RAG system. I’m super baffled as to what the safety instructions actually are other than “advanced anything” and even then their definition of advanced is a joke, I’m an idiot and my questions are almost laughable.
xscott 11 minutes ago [-]
Same here - my questions were sophomore level. I think it's notable that when I edited my question to say it was about Gemma 4, it answered without blocking. A cynic like myself would interpret that as evidence they don't care about sharing information if it involves their competitors.
runtime_lens 8 hours ago [-]
The dangerous part isn't that a model refuses extreme requests. It's when mundane requests become unpredictable enough that you stop trusting the model.
eli_gottlieb 12 hours ago [-]
Have they made Sol do less unwanted autonomy than the previous Codex models did?
hatthew 13 hours ago [-]
I feel like those examples are considered difficult because they're niche topics, but aren't actually all that difficult in a general sense. What I consider truly difficult are things like taking a ticket and implementing it in a preexisting codebase, using a clean and reasonable design that fits the existing style and makes sense to a human, and avoids the footguns I learned by working with the codebase for over a day.
jchw 11 hours ago [-]
If you said this in 2025 I would've 100% understood, but to be honest getting AI models to do a pretty good job on day-to-day ticket work has become so boring that we don't even bother using the top tier models and higher effort slots for that anymore. I personally wind up tweaking the results a lot and recursively having fresh agents review the diff, but that's just because I'm picky; in a lot of cases the first diff is actually pretty damn decent.
Compared to what I am doing at home experimentally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.
com2kid 9 hours ago [-]
All non-trivial code terra has generated for me has had at least one serious bug in it. Typically caught by a review from myself or Sol.
But I wouldn't trust lower tier models for end to end solutions.
jchw 10 minutes ago [-]
Personally I wouldn't want bots running autonomously on a repo, even if there were other bots cross checking them. But that having been said, I'd also say that my experience was similar with human code: it is rare to not find at least some issue worth at least pointing out. The only real difference is that the LLMs have vastly different holes than people do, making the real challenge trying to make sure you're covering them. In most cases for me the best solution seems to be just giving them a way to test and attempt to prove things out in a realistic environment. But clearly, we haven't really left the era of having humans in the loop. Fully vibe-coded codebases clearly suffer from a myriad of issues.
user43928 10 hours ago [-]
This is true in some sense.
Getting the AI to output code that you like is difficult.
As an example, let's say in React you have a "useLocale()" hook.
The AI will happily pass down locale as a prop to 5 child components instead of just calling the hook in the component.
A review from another model did not flag such stylistic issues either.
I believe that the latest models are very good at functionally achieving the goal, but still have poor taste for UX or code quality.
The most productive use of AI for software development happens in an environment where you do not review the code but test the UX end to end.
boorang 9 hours ago [-]
I use the AGENTS.md to show it how i want the code to look like. Something like "when implementing hooks adhere to the guidelines in docs/react-hooks.md". And then react-hooks describes your heuristics and what you consider best practices. There is a clear difference in code quality for me when using codex with a well crafted AGENTS.md vs. without one, you can run the experiment yourself pretty easily. As I mentioned in another comment, I think Claude poisoned users to stop relying on their Claude.md files and new codex users might be surprised at how well it adheres to guidelines.
user43928 8 hours ago [-]
We use skills for similar purposes, with positive and negative examples.
I think it sometimes worked, for example for testing preferences, but sometimes it did not.
Could be a problem with the harness also.
In any case, I feel that it's a bit playing whac-a-mole with explicit rules for things that a more intelligent model should do by default.
lightbendover 2 hours ago [-]
[dead]
edg5000 14 hours ago [-]
FYI I run it consistently in xhigh regardless of difficulty of the task at hand. I remember high being very fast, but I'd rather wait a bit more and get better output. AIs are insanely fast compared to me anyway, even on xhigh. Consumes more usage, but even at 100 EUR/m I don't hit limits.
saghm 14 hours ago [-]
After hitting the session limit on my company's plan so many times with Claude when I was using it, I mostly keep Codex on "high" rather than "xhigh" as a way to leave the tokens for my more ambitious coworkers. It's possible that having it higher might end up with better output, but so far at least I've yet to see a way to get any model to do 100% of what I need up front without any need for me to make changes that end up being more tedious to do via interaction than by hand, and it doesn't feel worth spending a bunch more tokens trying to figure out how to better communicate to it up front how the dominoes get set up so they fall in place properly the next time.
jchw 14 hours ago [-]
To be fair, I actually do run xhigh as my default. However, for the first time in my experience of trying and using LLMs, with Sol.. sometimes I feel confident enough to set the effort level to "Low". I just had Sol prototype some AWS stuff on low earlier. Great result, did exactly what I wanted.
FooBarWidget 10 hours ago [-]
How do you handle context limits? With more thinking tokens you fill it up earlier. Compaction degrades performance too. What's your strategy?
edg5000 7 hours ago [-]
Initially I was planning heavily around context limits, but I've learned to just ignore it completely. Compaction is seamless for me. If details are lost in compaction, the model just re-reads what's needed. My conclusion is that at least for Sol, the summaries (which I've never seen) must be amazing. Every now and then a detail gets lost and I have to repeat it. I don't think there is performance degration, because the model is smart enough to re-read relevant files as needed.
cgannett 14 hours ago [-]
And Mai-Code-1.1-Flash seems like a really good cooperative player to GPT 5.6 Sol. You get Sol to help you make a detailed plan, and Mai codes it up and you can get pretty decent code out the other end without too many tokens if you are careful.
manmal 12 hours ago [-]
Why wouldn’t you use Luna for that? It’s super cheap.
locknitpicker 8 hours ago [-]
> Why wouldn’t you use Luna for that? It’s super cheap.
MAI also offers a ultra cheap version that's competitive with Luna.
So much so that the models look like they were designed by a product manager explicitly to eat away OpenAI's market share.
Vscode even pushed them quite hard onto users with the latest release, going to the extent of putting up a modal to convince users to try them out.
alex7o 10 hours ago [-]
Taste I suppose?
swingboy 7 hours ago [-]
How detailed of a plan? Are you including code snippets or just behavior and letting the lesser model decide how to implement?
calvinmorrison 14 hours ago [-]
things that are alchemical are rarely alchemy. That is to say things are very fiddly but stick a room of monkeys on typewriters, a schizophrenic developer with HolyC and adderall or an LLM, persistence is the key to many of these things like drivers, extracting keys from vintage security domains, etc. Dropping into xdd to a human is a chore, not for an LLM.
jchw 14 hours ago [-]
Although I am not exactly sure what you mean, I am not really claiming it is doing anything I couldn't do - but yes, it does so with much less effort. For example, I can have it set up probes and tracing on Linux that I personally would have to consult documentation to do. It might not even have to consult the documentation due to having the information on-tap, but even if it does, it's nothing that would cause it any fatigue, it's just going to keep moving forward in a loop until it is satisfied that it meets the criteria. I could've done all of this alone - I really could have. I just would not have. Being able to do something 10 times faster or with 10 times less effort is, in some senses, sometimes more impactful than being able to do entirely new things you couldn't do before.
saghm 14 hours ago [-]
This is a really good point that I definitely failed to grasp when first hearing about these tools. At least for me, the best way to use these tools is as a way to free myself from having to spend time thinking about the things that aren't worthwhile so I can focus on the things that truly are. I've had times in my life spending hours reading documentation and googling random things to try to tease out the correct sequence of commands or the exact right shape of an API to be able to make things work to know that it doesn't make me more productive to do that myself rather than point an LLM at the thing and let it spit out the answer after a few minutes. Meanwhile, I can spend that time thinking about what comes next, or what the correct way to take that one-off output and abstract it to something that can be used meaningfully in more flexible ways.
The only obvious objection I can think of to this line of thinking (at least from a technical perspective) is "how does someone build up the knowledge to be able to use a tool effectively in that way if not by doing things by hand at first?" The honest answer that is "I don't know, but that's also pretty much exactly the type of thing my employers have never been paying me to solve in the first place". Even just a decade into my career, there have already been plenty of times in my career I've struggle to convince people that we should do stuff in a way that won't bite us in the ass a month or two down the line, and in the times I've managed to succeed, it's usually only by putting in more of my own time and effort to make the initial investment seem more palatable. Luckily right now I'm not in one of those times when I'm having to go full throttle to keep the lights on a few months from now, but I don't have enough fuel in reserves to work on a plan for when we need to build a new rocket in another ten years. Maybe ask me next month.
Der_Einzige 11 minutes ago [-]
Terry A Davis wouldn't have touched LLMs. He'd probably claim they're demonic.
wonnage 10 hours ago [-]
AI-pilled obsession with "taste" is bordering on insanity
It's just vibes
jchw 4 hours ago [-]
It's easy to dismiss "taste" when you either have none or just fail to appreciate it, but nothing gives you an appreciation for the importance of taste like LLMs. There is no benchmark for taste, so while many things improve taste does not. Bad taste is, in fact, a huge component of what makes AI slop so sloppy.
But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.
wonnage 15 minutes ago [-]
Thinking the giant array of GPUs has “taste” is exactly the problem
perching_aix 5 hours ago [-]
That's what they're always going to be, so not sure what would be "insane" about it. They literally feed on and emit natural language, and are put to work on informally defined, arbitrary tasks.
When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ, academic tests). This is not easy.
wonnage 11 minutes ago [-]
This sounds about as unhinged as Google saying Material Design 3 is 30% more rebellious
mfru 10 hours ago [-]
It's vibes all the way down.
It's... really just vibes?
Always has been.
kyxsc 14 hours ago [-]
Sol is way too eager to hone in on small details and ends up with massive over-engineering. Fable does it too - to be fair - but noticeably less.
After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.
Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value by far and there's no contest.
---
Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.
Both models suffer from doing-too-much. But both models are fundamentally really smart and knowledgeable. I think it's really close and pricing cuts really spice things up for us consumers! Sol is a clear winner in the value department and the $100 plan is enticing!
---
*Claude Code usage is reducing by 33% in 2 days, Wednesday August 19... cmon anthropic: clau.de/cc-50-promo
mwigdahl 2 hours ago [-]
I get a lot of mileage using Fable to spec, then Sol to review the spec, then Fable to plan, then Sol to review the plan, then Fable to implement, then Sol to review the implementation. It's a lot of steps, but a great boost in quality of output.
andreygubarev 9 hours ago [-]
Yeah, I very much agree on this. I think Sol and Fable code quality is on par. Maybe Fable is just a tiny bit better, but Sol compensates with its ability to work through things, while Fable, in my experience, generally tends to avoid solving problems that require many LOC.
However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).
One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.
11 hours ago [-]
greenavocado 14 hours ago [-]
Sol w/ Effort -> Low
kyxsc 14 hours ago [-]
It's great, don't get me wrong, but so is Fable. I'm just comparing the long-horizon task performance between the two at the same or similar effort levels.
Given the 50% discount on Sol and how smart it is, yeah it's unprecedented value. If you only want to use low effort, there's a clear winner here on value and it's not even close!
iJohnDoe 14 hours ago [-]
Interesting the use of Max and Ultra. I don’t doubt the complexity, but would someone use Max or Ultra on Typescript or Go, for example?
Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?
kyxsc 14 hours ago [-]
*The "Max" I referred to was the plan tier, not the effort level btw
For small tasks, you can just use something like low or medium effort and it can usually avoid mistakes; after all, the model will test the code anyways and can do some baseline level of iterating.
In regards to cost, we need to acknowledge how generous OpenAI was in the last couple months with Codex usage credits (no weekly limits) and usage resets. It afforded me many a dive with Codex! Yes it uses more tokens, but sometimes it's worth it -- just depends on what you're working on.
Finally, Ultra(code) isn't that bad when it comes to cached tokens. I think folks overstate the general token usage of ultra effort on both providers.
---
Both models are great at green-fielding a project when given detailed specs.
Both models overthink too liberally (imo) during these larger multi-shots. Sol overthinks more than Fable.
Both models are really smart and perform great for general knowledge and regular coding tasks.
FL410 15 hours ago [-]
Fable feels less cumbersome to work with, but it is SO DAMN ANNOYING with the refusals that I'm leaning more and more on Sol, and very much looking forward to GPT6. Just seems like Anthropic is trying their hardest to ruin their reputation and user experience.
paxys 14 hours ago [-]
Remember, Dario knows best
teaearlgraycold 14 hours ago [-]
I’ve never had it refuse anything. Even vulnerability searching in my codebase.
ChadNauseam 10 hours ago [-]
Yeah I feel like I'm living in a different dimension than these people. I wonder what they're working on. I've literally never had it refuse everything and I max out my 20x plan every week
SyneRyder 9 hours ago [-]
Can I ask which country you're in? I have a theory that the safeguards differ depending on the user's country.
I'm in Australia, and Fable downgrades to Opus when testing for bugs in memory in a legacy C code base. If Fable starts taking initiative and writes a test case that involves writing to a null pointer, that's the end of the conversation.
FL410 9 hours ago [-]
As one random example, today I had it hit a refusal loop when adding a country selector dropdown to a form, presumably because it contained a “bad” country name? I hit refusals at least 2-3 times per day, sometimes many more. The worst part is it is often right in the middle of a multi-stage task, so the only option is really to switch to Opus 5 and let it defecate its absurdly verbose comments all over the rest of the edits in the turn and hope it doesn’t go on one of its tangents, then have Sol do damage control. Oh and I got approved for their “cyber verification program” blessing, which comically does absolutely nothing for Fable.
leokennis 9 hours ago [-]
5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more/less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.
matheusmoreira 11 hours ago [-]
I too switched to OpenAI after I got sick of Anthropic's constant "safety" downgrades. Sol is definitely a breath of fresh air.
> even after completing the verification program
Was it easy to complete it?
I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.
Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.
ec109685 15 hours ago [-]
Fable is still the best there is. Sol close second but I find it gets way to stuck on details.
Also Opus 5 is fine if your codebase is simple.
curreylabs 14 hours ago [-]
And you arent writing English
boorang 10 hours ago [-]
I was pleasantly surprised to find that the GPT models are much stricter in adhering to my AGENTS.md guidelines and heuristics than Claude.
Aargau 15 hours ago [-]
I'm also in Anthropic Cyber Verification Program, but they specifically exclude Fable, just goes up to Opus 5.
I hear you on the downgrades, I'm 13/13 on downgrades, and last downgraded me to Sonnet for asking for reasoning chain.
justinbaker84 5 hours ago [-]
I switched from claude to gpt when 5.6 came out for the same reasons. I don't understand why so many people are still using claude when GPT and open source are so much better.
znnajdla 11 hours ago [-]
Sol is my daily driver but there are still times I reach for Fable when Sol doesn’t cut it. Just yesterday for example, I was trying to build a self-modifying hot-reloaded agent harness in Elixir for fun and Sol just kept doing silly things like thin wrappers and unnecessary abstractions. Fable handled the task elegantly. Sol is really good as a reviewer for finding bugs due to its thoroughness however.
shibaprasadb 8 hours ago [-]
I cancelled my subscription recently and moved to Sol. So far - it has been a great experience. The only aspect where Fable/Claude is better I feel is doing some research from the web and summarising the facts.
nutribeatApp 6 hours ago [-]
I generally prefer Fable but in my experience Sol is a much better web researcher
jamesponddotco 3 hours ago [-]
Fable is the only one that follows my instructions correctly, which I find quite important. It takes STYLE.md and SPEC.md as law, and code just like I would code, with the same mistakes and all.
I just can't get Opus (Opus 5 is dumb as a rock, to be fair) or Sol to do that, so I exclusively use Fable for personal work. When I reach my weekly limit, usually on the last day close to the reset, I just go back to coding by hand ¯\_(ツ)_/¯
Heck, it even does security reviews and fixes, as long as I don't ask it to "attack" the codebase. I'm planning on using Kimi or GLM for that part.
impulser_ 8 hours ago [-]
It's the complete opposite for me. The model might be the worst model I have ever used when compared to other models in the class. You just can't get it not to just write the most enterprise over complex over engineered solutions for every little thing you ask it to do.
It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.
I don't even want to see the codebases this model is fucking up.
It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.
tamimio 14 hours ago [-]
Yeah at this point claude is overrated, overly expensive, weird writing style (elliptical), and the worst part is the aggressive guardrails that even normal convos get interrupted, meanwhile openAI is still I would say at the normal balance, if you ask something too obvious or direct it will stop you other than that, it work flawlessly, plus, I have yet to hit the limit despite heavily using it these past weeks.
ChadNauseam 10 hours ago [-]
> if you ask something too obvious or direct it will stop you other than that, it work flawlessly
What on earth are you asking it?
bbg2401 4 hours ago [-]
You've replied incredulously to a similar stated experience in this thread already and proceeded to ignore the follow-up. Why are you again asking a question to which you have no intention to field an answer?
synergy20 4 hours ago [-]
indeed, I'm switching from claude to codex
kingkongjaffa 3 hours ago [-]
I also cancelled Claude recently. GPT 5.6 models are really good, and they have much better usage limits.
johnnyApplePRNG 5 hours ago [-]
Really?
I have witnessed 5.6 Sol Ultra edit line after line of literally empty lines ... for hours.
I wasn't literally watching it, I came back to a goal (that it started for itself without my approval!) that had done nothing but that for some reason.
It couldn't explain why it had started.
vdntp 10 hours ago [-]
you should check out the codex desktop app. people who've been using claude code for a long time will surely be surprised.
jiaosdjf 8 hours ago [-]
The final straw for Claude was its refusal to give me a list of the most recent rapes reported by the BBC and basic information about them (location, date, names, just things reported in mainstream media). It outright REFUSED to complete this task.
I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.
arcrs 6 hours ago [-]
output token efficiency bruv. u never go wrong with it
saidnooneever 9 hours ago [-]
sol is much better imho than Fable but i can understand if they will perform wildly different for different people with different levels of expertise aswell as different needs. I dislike fable myself it doesnt really work for me.
Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.
cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).
what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...
for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.
nurettin 6 hours ago [-]
Used claude since 7/2025. Switched to codex after fable got blocked. It was still 5.5 but I knew they had to come up with something. As soon as I switched, wow. It wasn't super intelligent, but it was stable. Every day it was the same performance. This consistency is definitely worth paying for.
yieldcrv 11 hours ago [-]
Claude as a harness at all really spends too much time before giving user feedback
Its a crutch that is no longer competitive
petesergeant 11 hours ago [-]
I have a "strategy / life-coach" project, and was surprised at how much better Sol is than Fable on it, as I've found Fable to have the edge for most things for me so far. But Sol: questions were better, insight was better, it got the brief better.
Kye 13 hours ago [-]
Sol has held stuff for a while to do the same sort of hazard checks I assume Fable is doing, but it always releases them. I think that's the better way to handle it rather than preventing me from seeing how far I can get generating schematics to use in Minecraft. Currently: a mostly normal voxel house.
Razengan 15 hours ago [-]
I recently tried Claude again after several months, to see if it was any better at something Codex has been struggling with…
They STILL don't have an option to "Sign in with Apple" on the website, but they do for Google??!? (and on iPhone of course)
Screw that asinine UX
(and no it wasn't better than Codex at this particular task)
nozzlegear 11 hours ago [-]
That was an issue at least a year ago. I had signed up for a claude account on my iPhone and then wanted to sign in on my laptop but nope, not possible. Insane they still haven't fixed it.
Can somebody at Anthropic tag claude in slack or whatever goofy shit you do and ask it to add Apple OAuth to your website? Clearly humans aren't testing it.
Razengan 7 hours ago [-]
I signed up on iOS, Sign In with Apple, cause I don't go around giving random companies my actual email if I can help it
and sure enough, I was right to do so: They don't even let you remove your payment method afterwards. Every other store, Steam etc., lets you.
No way I have enough trust to install their desktop app after that, so I just want to try it through their website..
Can Sign In with Google, but not with Apple
so you gotta open the Passwords app, copy your random email, paste into the website, then copy the OTP from your email..
It's been that way for at least a year
The desktop app was clunky too the last couple times I tried it a few months ago
So all the Claude hype posted on HN seems like a case of the emperor with no clothes to me
(P.S. The thing I just now tried to do on Claude hit the weekly usage limit after 2 minutes)
solenoid0937 6 hours ago [-]
There is no Claude hype on HN, HN has more vitriol for Claude than any online forum I've seen
datakan 5 hours ago [-]
Can't change your email address with Claude either. You have to delete your whole account and create a new one. Absolute shit software and they expect us to believe these things are super powerful world changing things.
onlyrealcuzzo 12 hours ago [-]
This sure looks like a race to the bottom to me, and I love it.
If Sol isn't the best model, it is up there...
You don't cut the price of the best model for no reason...
dvt 10 hours ago [-]
> This sure looks like a race to the bottom
Always has been. My prediction is that both OpenAI and Claude will go bust unless they deliver a killer product. And unlike scrappy startups, they have a pretty serious deadline because creditors will come a-knockin'.
There's little to no functional difference between Kimi, Qwen, Sol, Opus, etc. All flagship models are within like 1-5% of each other and the real moat will be what's always been the hard part: making a good product.
dcre 3 hours ago [-]
It's hard for me to imagine how you could define "killer product" to exclude something that takes you from $9B to $65B ARR in 8 months.
The Chinese models are cheap because no one is using them. But they can't actually afford (or have capacity) to serve enough people to kill the giants. This is evidenced by them all recently hiking prices or limiting usage.
It's possible that they build out in China at an unreal pace, China doesn't have concept of "community input" to drag down state projects, but then you are left giving your IP to China. Just ask western hardware businesses how well that goes.
notatoad 2 hours ago [-]
Isn’t almost all of anthropic and OpenAI’s compute leased?
If compute is the moat, that doesn’t really make their position any less precarious.
Der_Einzige 4 minutes ago [-]
No one is in a better position to more efficiently use it - that's what happens when you poach every top 0.01% engineer/researcher in AI.
They're guaranteed to get over whatever hump you think they're in unironically. Uber/Tesla have been in far worse situations and despite Elon being an idiot/liar you see how they performed when even the most bullish of investors called for their heads
matheusmoreira 9 hours ago [-]
> All flagship models are within like 1-5% of each other
Don't know about that.
I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc.
Ran a complete Fable/max code review. Took over a month on a subscription. Now I've switched to OpenAI and am repeating the exact same review with Sol/max.
It's still not done yet but preliminary findings suggest Sol can only reproduce 70-90% of Fable's findings. So I think these models aren't as close as we've been led to believe.
beacon294 8 hours ago [-]
Check the remainder for hallucination for sure.
unreal6 1 hours ago [-]
check all for verifiability
huflungdung 9 hours ago [-]
[dead]
aronowb14 2 hours ago [-]
Have you found an alternative to codex / Claude code? I’ve tried opencode but it’s been buggier / feels like an inferior product to both.
blfr 8 hours ago [-]
There is a massive difference even between Opus and Fable, same provider, before various harnesses and other optimizations come into play. Don't be deceived by rankings and benchmarks, try for yourself.
roncesvalles 8 hours ago [-]
The problem is that most of the volume doesn't come from proprietary products, it comes from API use which has no stickiness.
Claude already has a killer product (claude.ai/chat is a Swiss army knife) but just relying on people typing stuff into chat is not enough to sustain the company.
The other strategy is entrenching yourself as the LLM of choice into existing products (like ChatGPT is on Apple products).
sumedh 7 hours ago [-]
> All flagship models are within like 1-5% of each other
Depends on your use case. the Chinese models are not there yet.
But who can take a benchmarking website seriously when they literally change their benchmarks to appease anthropic, like artificialanalysis.ai did the other day?
andai 12 hours ago [-]
You do it if you can afford to do it and your competitor can't.
psadri 12 hours ago [-]
Who are OpenRouter’s competitors?
aurareturn 10 hours ago [-]
OpenRouter doesn't decide on the pricing.
resonious 10 hours ago [-]
They do decide on the 5% markup. But as far as I can tell, all other routers just match 5%. Not sure what they're competing on.
aurareturn 9 hours ago [-]
Yes but Sol dropping in price by 50% is not OpenRouter deciding. It's OpenAI.
mcintyre1994 9 hours ago [-]
I don’t think that’s true. OpenAI docs don’t have this price change. I assume they’d be the source for this post if it was true. The banner on OpenRouter for me says Gemini 3.7 discounted for a limited time, but if I click through that I get to this page: https://openrouter.ai/models?discount=true
That shows a bunch of models, including Sol, with a discount. None of them say how long it’s for, but I’d assume in all their cases it’s for a limited time as the banner said, and only on OpenRouter.
aurareturn 9 hours ago [-]
OpenRouter margins are not 50%.
stanac 9 hours ago [-]
As I understand OAI is offering discount only for users using the model via OR. Not sure why. Maybe they want OR users to try the model and switch to OAI subscription or something.
SyneRyder 6 hours ago [-]
It might also be worth noting that the discount is only if OpenAI itself is used as the provider via OpenRouter. The discount does not exist for Azure or Amazon hosted Sol.
I posted elsewhere, but the Azure uptime & performance for Sol is truly dire. OpenAI is offering 5x faster latency, 4x faster tokens generation, and vastly better uptime (Azure US has only 87% uptime), all for 50% of the price now. I assume the pricing is to compete with other shiny new models (Grok, Qwen etc), but it might also be to cut-off a truly poorly performing Microsoft hosting experience.
aurareturn 8 hours ago [-]
Yes, but this discount must be coming from OpenAI and not Open Router.
fileeditview 11 hours ago [-]
That kind of is a reason.
livinglist 10 hours ago [-]
The lower the merrier.
oblio 9 hours ago [-]
Competition is good for users.
livinglist 9 hours ago [-]
Sorry I meant to say lower, price that is, corrected.
Fergusonb 17 hours ago [-]
Luna saw a huge jump after the price cut and is one of the more competitive models at the new price on openrouter.
Maybe they want to see how much market they can grab with Sol?
This might help but there are already cheaper models with Sol's intelligence more or less, the most notable being Grok 4.6 at $6/m which makes it a tougher sell
xmonkee 17 hours ago [-]
It's really only between Anthropic and OpenAI for many of my use cases, since I have a Zero Data Retention agreement with both. I'm not trusting random inference providers and especially not Elmo with sensitive data.
speedstyle 14 hours ago [-]
or Tinfoil [0]? They serve open models with container integrity attested by Nvidia/AMD enclaves. Every cloud provider offers this of course, but not usually in a way that can be shared between distrusting users for economical inference. It still relies on the open-source containers being secure, and there's probably hardware sidechannels and stuff, but personally (ie privacy not liability) I trust it more than a contract
tinfoil asks more than 10x the output cost ... $1.90 per 1M tokens instead of $0.18 per 1M tokens for my favorite model (Deepseek V4 Flash 0731) on my favorite provider (DeepInfra) currently, for example.
johnnyApplePRNG 3 hours ago [-]
Deepinfra is US based and has ZDR FYI.
I am quite happy with them so far.
c0rruptbytes 15 hours ago [-]
plenty of companies offer ZDR and are just as random as OpenAI and Anthropic in their age
xmonkee 14 hours ago [-]
Would love a few names. Many of them fail to provide good uptime for large scale jobs or don’t have batch apis at all in my research.
HDBaseT 12 hours ago [-]
DigitalOcean does inference, has good uptime and offers ZDR. They are not an AI first/inference first company. They have been serving cloud products for almost 2 decades.
Batch API is a bit harder, not many models/providers support Batch. It primarily is only the Gemini/ChatGPT/Claude models that do. DigitalOcean does support a 50% discount on Batch API via them directly, not listed on OpenRouter.
For Mythos and even Fable they require prompt retention on their end.
edit: or more precisely if you want to access Mythos/Fable ZDR does not apply, and depending on config the exclusion can affect other models.
mrngld 5 hours ago [-]
I don't have experience directly with Anthropic, but I wouldn't assume anything with such high confidence without knowing the persons situation. No, an individual off the street or with an LLC and 5 employees isn't going to get a special deal from someone like Anthropic.
But if their employer is bringing millions of dollars of potential spend to the table, can tell you from years of experience that turns a lot of 'no's' to 'yes'.
robbru 3 hours ago [-]
I use Notion Enterprise specifically for its Zero Data Retention (ZDR) access. Fable does not offer ZDR, so Notion provides a dedicated settings toggle that disables Fable outright.
kasey_junk 5 hours ago [-]
I work for a large enterprise and fable/mythos are specifically banned because of the difference in data retention agreements from those models to other Anthropic models.
joshheitzman 14 hours ago [-]
[dead]
101011 13 hours ago [-]
There's a huge jump between OpenAI, Anthropic, Google, and every other major player distilling the internet into LLMs and deliberately breaking a mutually signed contract between them and another business.
As for ZDR and court-orders, what would you rather happen there? Violate the law or comply with holding the data? I would bet that any ZDR agreement has this court-ordered risk mutually understood and agreed upon.
I'd be surprised if they closed out over 100B tokens today (they did 101B on sol 5.6 yesterday).
OutOfHere 17 hours ago [-]
Since when does Grok 4.6 have Sol 5.6's intelligence? I don't believe it.
maxdo 15 hours ago [-]
I do have free sol and cursor ultra for 200 I prefer grok over sol, they are equally capable but grok is faster
qingcharles 11 hours ago [-]
I use Grok 4.6 every day; it's good, but it's not Opus 5 or Sol 5.6. The gap is closing, though.
dimgl 16 hours ago [-]
Why not?
chaos_emergent 16 hours ago [-]
Because it’s good on benchmarks but not on real usage?
dimgl 16 hours ago [-]
But OP said they've never used it. How would they know?
redox99 15 hours ago [-]
It doesn't.
mohamedkoubaa 17 hours ago [-]
I wonder if xAI is A/B testing routing some difficult grok 4.6 queries to Sol to seed some true believers.
jLaForest 16 hours ago [-]
[flagged]
bko 16 hours ago [-]
[flagged]
kelvinjps10 10 hours ago [-]
I have switched to Chagpt sub now after only using Claude for coding. You get more value for your money and feels like codex has reached Claude code performance in coding (the reason for using Claude) regular plus account allows you to have access to their most powerful model, image generation and asking questions is better because you can use sol but in instant mode and it feels smarter and faster.
And finally codex usage limits are better than the Claude daily 5h limit.
And codex feels faster although Claude code had more features
gb2d_hn 10 hours ago [-]
I switched as I felt Codex was on a par with Opus, but the chat responses from Sol are just more intelligible than the word soup I've been getting from Opus. I wonder if Opus could be prompted to respond in simpler prose via agents.md
no_no_no_yes 25 minutes ago [-]
Yes it can, it's called "output style" under `/config`:
```
This changes how Claude Code communicates with you
1. Default Claude completes coding tasks efficiently and provides concise responses
2. Proactive Claude executes immediately, minimizes interruptions, and prefers action over planning
3. Explanatory Claude explains its implementation choices and codebase patterns
4. Learning Claude pauses and asks you to write small pieces of code for hands-on practice
```
You can set up a custom one under `vi ~/.claude/output-styles/eli5.md` and `eli5` would show up in the list above:
I can't process lengthy text, talk to me like I'm 5.
Small words, short sentences, short paragraphs. If you have to use a big word, explain it right after. Only return what's actually necessary. Just tell me what you did, did it work, what do I do now.
If I have to decide something: show 3 options max, the context I need to pick fast, and which one you'd go with.
Keep paths and commands exact. I have no brain cells left for the rest.
```
However, the fact that you can do this says something in my opinion. And it doesn't necessarily point to somewhere good if I'm being honest. These niche config options are kinda crazy? I've been using CC for the past year and just learned about this last week, I don't know if it's new, or whether it's been there a while, but I feel like it shouldn't be necessary.
FluffyPancake 9 hours ago [-]
I have also noticed that I am increasingly struggling to read what LLMs are writing, finding it incomprehensible half the time.
I stole Matt Pocock's line of "When reporting information to me, be extremely concise and sacrifice grammar for the sake of concision." for the agents.md
It makes it a little better. I also specify to use https://github.com/AminBlg/SimpleEnglish/ for all writing it does including code comments. It all feels like a bandaids but that seems to be the best we can do right now.
marcyb5st 9 hours ago [-]
Ah, so I am not the only one struggling with deciphering Opus writing style. At times I feel dumb as a rock because I read the same passage like 5 times and I still don't get it.
jen729w 9 hours ago [-]
> I wonder if Opus could be prompted to respond in simpler prose via agents.md
God knows I've tried. I've got a variant of the ASD-STE100 trick which does the job, mostly, at the start … but get to about 100k of context and it goes out the window.
The model's personality is too strong for simple suggestion, alas.
andypants 6 hours ago [-]
> codex has reached Claude code performance in coding
Codex has always beaten claude in coding benchmarks, hasn't it?
CompoundEyes 16 hours ago [-]
I used over a billion tokens per day of gpt-5.6 sol xhigh starting last Wednesday through Sunday before reaching my reset limit. The $200 pro plan is still the best deal.
ec109685 15 hours ago [-]
I spent $800 in a few hours when my sub maxed out because I was trying to get something done and had a long car ride to let it churn.
Their api pricing is absurdly expensive.
giancarlostoro 15 hours ago [-]
> Their api pricing is absurdly expensive.
I assume at this point that it subsidizes subscriptions.
timClicks 14 hours ago [-]
Subscriptions are a mechanism to attract developers, who then advocate that their company should use the API.
HeWhoLurksLate 14 hours ago [-]
yes, it absolutely does.
I've gotten more work done on a second chatgpt pro $100/mo subscription than I did with ~$150 of paying for usage through the app.
qingcharles 11 hours ago [-]
Massively subsidized. As soon as my Claude switches from subscription to overage I have to tap out quickly.
skohan 10 hours ago [-]
This is a big part of the reason I went local-only. Subscription limits are horrible for having a decent workflow.
LinXitoW 2 hours ago [-]
There's no universe in which just buying a second or a larger subscription isn't a billion times better and cheaper than any kind of comparable local workflow.
Privacy, experimenting with ML and "unorthodox" needs are currently the only acceptable reasons to do local.
skohan 37 minutes ago [-]
Well I wouldn't say a billion times better. I've actually been having a surprising amount of success working with local models. And my investment has only been the equivalent of 4 months of a Max x20 subscription.
Experimentation and privacy are definitely advantages, but it's also quite a lot of fun.
giancarlostoro 5 hours ago [-]
Yeah, it sure was convenient that there was a RAM pricing crisis right when Apple was making local inference viable. All because of a promise that AI companies will buy more of it... with money they don't yet have, whereas Apple does have lots of money.
CompoundEyes 4 hours ago [-]
Agreed I’ve seen what they’re paying at work for the OpenAI API but I also think that includes reserved capacity and ZDR O_o. Try some add on credits next time if you can. I was curious at how far $20 would go (500 credits). Watched them go to zero over an hour and assumed it’d stop. It then ran for another 6 hours and completed the task despite the meter at 0. What a task means is very unclear but it’s definitely not pricing sol at $20/500 credits an hr in tokens.
paxys 14 hours ago [-]
Because API pricing is for corporations and subscriptions are for consumers.
blitzar 9 hours ago [-]
I thought corporations were meant to be smart and get b2b / volume discounts - not pay 5-10x what the man in the street is paying.
oblio 9 hours ago [-]
FOMO. This won't last forever.
wmf 12 hours ago [-]
Are we supposed to just shrug at the idea of businesses paying 10x more for raw materials than consumers? How long can this go on?
zipy124 2 hours ago [-]
I mean this is kind of standard in a lot of industries? Businesses get reliability/support usually for the price increase. Just look at lots of industrial tech. Or business vs consumer 3D printers, or business vs consumer laptops etc....
kelvinjps10 10 hours ago [-]
Hopefully for a long time this is like the beginnings of a vac startup, enjoy while you can before the enshitification comes
SJMG 15 hours ago [-]
A billion a day? How many agents are you running?
_zoltan_ 9 hours ago [-]
with ultracode, it goes fast. I can easily get to a billion on a busy day.
greenavocado 14 hours ago [-]
That's wild. I was gonna say 3-5 billion a month is more reasonable summed across all token types.
XCSme 7 hours ago [-]
I am also on the pro $200 plan, the limit is high enough to do what I want for a week, and binge run Ultra Fast last day to use the remaining credits.
infinite_spin 16 hours ago [-]
I have mine churning like butter and I'm rarely hitting a billion tokens per day, what's your workflow look like?
CompoundEyes 15 hours ago [-]
Pretty basic. The codex app with one conversation per project and several running simultaneously all hours. I’m going for max caching that way and it never gets lost even with compaction somehow. Each has a plan with milestones to keep up to date and a thin agents file. I check in on them in the Remote app. Use case is protocol and control reverse engineering of audio hardware. I think they must be identifying the heavy use agent sessions and cranking up their cache lives so it’s not a big deal for them.
rpdillon 15 hours ago [-]
A billion tokens a day is 11,000 tokens a second sustained. How many tokens per second are you getting off of GPT 5.6 Sol per project?
sheepscreek 15 hours ago [-]
Ultra mode spins up many sub-agents. On a particularly challenging task, I’ve had as many as 29 agents working at one time.
Also if you don’t specify, most end up being the same as the parent model which is pretty wasteful.
I engineered a skill that spins up Terra High agents for most sub-agents, resorting to Sol Medium for technical research and Luna High for code/in-project research tasks.
On a slightly different topic, Luna Max is incredibly capable and doesn’t use as much quota (Luna tokens are dirt cheap).
iJohnDoe 14 hours ago [-]
Most importantly, are you seeing a return on investment for time and ultimate outcome?
No one can judge the enjoyment, learning, and hobby aspects. Just wondering if there is an end goal for that much overall expenditure (time, money, energy, etc.)
_zoltan_ 9 hours ago [-]
I absolutely see a return. I do not pay for this (we have a corporate gateway), for the price of a junior developer I can get 3-4-5 senior developer's work done. it's insane value.
jeffmcjunkin 14 hours ago [-]
Often people are counting all tokens, including cached input tokens, for those more impressive "billions of tokens" quotes.
rpdillon 13 hours ago [-]
Ah, thanks, I'd missed that nuance in the other reply!
childintime 6 hours ago [-]
I'd like to see a benchmark on this specific topic: Reverse engineer the hardware protocol from a driver, or just migrate a driver from one OS to another.
CompoundEyes 4 hours ago [-]
I’ve chipped on it with each model since 5.2 but 5.6 sol is something else. When it first came out I’d get some refusals but they’ve since stopped. I wonder what an ideal candidate benchmark task would be for that?
ramraj07 15 hours ago [-]
Some people just do crazy stuff. For example this now ex yc guy who said he has agents constantly scanning Sf govt apis and forming dashboards just because
jimbob45 15 hours ago [-]
Is anyone hitting caps without agents or API usage? Seems very difficult.
brookst 15 hours ago [-]
Yeah I am, building large software with a vision - requirements - architecture - plan - code workflow. One Claude max account is enough to work on one, maybe two of those at a time (call it 15B tokens/month per project)
dyauspitr 15 hours ago [-]
I’m exclusively using ultra and I run out in 3-4 days consistently. Those resets are great but I’ve noticed they like to cluster them at the start of the cycle, would be better if they spaced them out more.
IshKebab 14 hours ago [-]
A billion tokens per day?? Plausible estimates put the energy use at about 0.001 Wh/token, which means you're using 1000 kWh/day in electricity, just to generate slop. That's about the same as 50-100 houses. 300kg of CO2 per day - roughly the same as flying from London to New York every three days.
I think on average AI energy usage is not as big a deal as everyone is panicking about, but your usage is truly absurd and I don't know how you can live with that. It's immoral.
kolinko 8 hours ago [-]
That’s for i/o tokens, mostly output. 90-98% is cache read usually, so you can divide electricity use by 10 at least.
As for co2, it depends on the provider, it could be way lower as well.
As for ethics, you don’t know what he works on, and how effectively - he might be saving 10x that much of co2 for the planet.
spider-mario 2 hours ago [-]
Where do those estimates of 0.001 Wh/token come from?
Gangway0829 14 hours ago [-]
I can't speak for that guy, but I'm a physicist and work in clean energy... So it's not too hard! That said, I usually am closer to 10M on days I do heavy coding, so not nearly that bad.
breezybottom 4 hours ago [-]
Terrifying to think that even physicists can't code their own simulations anymore. We're plunging headfirst into the dark ages.
paxys 14 hours ago [-]
Token caching is a thing
throwup238 14 hours ago [-]
You really think OpenAI is selling $1500/mo of electricity (at $0.05/kwh) for $200/mo?
I’m guessing that Wh/token estimate is several orders of magnitude too high.
Noaidi 14 hours ago [-]
They certainly could be using that much electricity at a loss based on their profitability, which doesn’t exist.
Leaked financial documents from 2025 show the company reported an operating loss of approximately $20.9 billion against $13.1 billion in revenue.
blovescoffee 12 hours ago [-]
As a company... but that includes things like research costs, model training etc. to determine if they're selling electricity at a loss you should look at inference costs bc that's the "thing" they're selling
WASDx 8 hours ago [-]
I'm glad someone is voicing this. Overconsumption at that level is not defensible. However if they meant cached tokens so it's not that bad.
pertymcpert 11 hours ago [-]
Most tokens are cached.
lannisterstark 14 hours ago [-]
1. You do not know what they're using it for.
2. Get off your high horse please.
3. Immoral my ass.
nozzlegear 11 hours ago [-]
> Immoral my ass.
Is any amount of tokenmaxxing moral?
stavros 9 hours ago [-]
Is any form of energy usage moral?
nozzlegear 1 hours ago [-]
Only for plants, those weird bacteria that live near thermal vents, and maybe the tardigrade.
f6v 5 hours ago [-]
> just to generate slop
Do some people still deny you can do a shit ton of work with AI?
_zoltan_ 9 hours ago [-]
"slop"? come on. we're not in 2020 anymore, Dorothy.
jm4 16 hours ago [-]
I can’t sign up for that. I tried authorizing Codex a couple days ago. For some reason, their system says my phone number has been used for verification 3 times even though it definitely has not. I’ve had this phone number for over 20 years. OpenAI support is useless. They just keep repeating the policy without actually helping me.
weakened_malloc 14 hours ago [-]
Use TextVerified, load up like $5 of credit and OAI verification is like $1.00. Then when your account is made, ensure 2FA/passkey is setup then you don't need to worry about the phone number.
stavros 9 hours ago [-]
Thanks for that tip!
jaggederest 15 hours ago [-]
Get a burner and use it? If you're spending $200/mo on something, $40 or whatever for a burner phone seems like a pretty cheap price.
63stack 7 hours ago [-]
Oh no fuck that, business 101 is make sure that your checkout page works. There is plenty of competition in this sector, take your money elsewhere.
paxys 14 hours ago [-]
You can get a phone number online for a few dollars.
jaggederest 14 hours ago [-]
Historically those are less useful because some of the verification systems require a real phone number and that your name is associated with the account, depending on what and how they verify. It's annoying, I use a google voice number as my primary, and it often gets rejected.
qingcharles 11 hours ago [-]
Good2Go is $5/mo for a real SIM with unlimited talk/text + 1GB data.
_345 13 hours ago [-]
Yes I filed a support ticket with them and explained that their system is broken and they just did not care. I explained how it was impossible for me to use it 3 times already as I've only made 2 chatgpt accounts EVER, and only recalling entering my phone number for one of the two chatgpt accounts. I told them that this issue locked me out of codex and chatgpt for work and they weren't willing to do anything about it. Totally useless support.
I ended up borrowing my gf's phone number just so I could get access for work. Ridiculous
You're literally encouraging someone else to come in and steal your customer base,
ywvcbk 6 hours ago [-]
Or it's "normal" market segmentation. OpenRouter users are more price sensitive in general, also a lot of enterprise users who can't switch easily are using the official API (or Bedrock or Azure) and you want to squeeze them as much as you can.
christina97 2 hours ago [-]
They’re trying to compete with open weight models here. It’s a very different customer segment.
dannyw 6 hours ago [-]
A surprising amount of companies sell exact the same product through different channels for different prices. A good example is Apple.
Multiple times a year, retailers here in Australia have co-ordinated sales on Apple products. Apple.com or their retail stores don't have these sales.
But they're clearly Apple-funded when competing retailers launch the same sales on the same days; and the margins aren't enough for retailers to take a loss.
podgorniy 8 hours ago [-]
This imbalance of exchange means that OpenAI is getting something from this deal. Question what is exactly.
SyneRyder 7 hours ago [-]
Seems reasonably clear to me? Potentially bringing in more customers who use OpenRouter for trying out all the models with rapid switching, enticing them to use Sol. And anyone being routed on price will immediately be switched to OpenAI's servers instead.
It also seems to be providing a vastly better user experience - Azure has less than 99% uptime (Azure USA only has 87% uptime), latency of 20 - 30 seconds, and a mere 8 tokens per second. OpenAI is offering 32 tokens per second (4x faster), 4 seconds latency (5x faster), and all for half the price of what Microsoft is charging for a vastly inferior experience.
I think being able to A/B test price is invaluable for them. They can't cut prices by 50% and hike it again. By letting others slash the price, they can tell if it is worth it or not to do this officially.
rafaelmn 6 hours ago [-]
Literally all OEMs do this, if you go to Apple site you won't see any sales, but Amazon and other retailers will have stuff at 20% off regularly
oblio 9 hours ago [-]
That's weird.
JCharante 8 hours ago [-]
Maybe they (oai) want to pump their marketshare on openrouter lol
cute_boi 40 minutes ago [-]
Any benefit of pumping marketshare on open router?
not-kinsale-joe 9 hours ago [-]
Why is that weird?
Mattrou 8 hours ago [-]
How is that not weird?
Has OpenAI struck a deal with openrouter and that's why we're seeing preferred pricing?
Is openrouter taking a loss on sol API calls to grow adoption?
How temporary is the reduction in price?
oblio 8 hours ago [-]
Selling cheaper through what should be a minor third party, than through the first, party is super weird in commerce.
krzyk 7 hours ago [-]
So the title is misleading, intentionally.
z_rho_one 16 hours ago [-]
If they can cut the price of Sol by 50% and the price of Luna by 80%, then the original price might have carried a massive operating margin. They might still be serving the models at a profit after these price cuts, but we will never know.
paxys 16 hours ago [-]
I don’t think there’s a real answer for this. Margin depends on whatever number the accounting department wants to make up.
Do you include research and training costs? Of all models or only the ones being served? What percent of the R&D budget do you allocate to inference? What about data center capacity? Do you count future commitments? All the circular financing deals? Do you count employee equity grants as costs? At what valuation?
dannyw 6 hours ago [-]
We have a simple definition for this: COGS.
We also have another solution for "whatever accounting decides": generally accepted accounting practices. It's far from perfect, but GAAP figures are what you should be looking at; not "adjusted GAAP" or whatever invention.
wahnfrieden 16 hours ago [-]
OpenAI didn't cut the price of Sol by 50% like they did with Luna's 80%. Sol was unchanged. This is just a limited promo for OpenRouter non-BYOK.
kolinko 8 hours ago [-]
Or they have gotten new asics and can do now inference way cheaper
XCSme 7 hours ago [-]
Why would they infer faster with new shoes?
tclancy 6 hours ago [-]
It is a well-accepted fact new shoes make you faster. Current science suggests it is due to the lighter weight from lack of dirt, though there is a competing theory which says it's an optical illusion due to the fact the pure white streaks resemble speedforce.
InsideOutSanta 10 hours ago [-]
I'm pretty sure tokens are priced to maximize revenue, not inference profit.
josu 10 hours ago [-]
I always find it funny that Japanese pensioners are probably subsidizing my tokens.
dgellow 8 hours ago [-]
What do you mean?
josu 7 hours ago [-]
SoftBank is one of the largest investors in openAI having contributed more than 30B and pledged another 30.
dgellow 5 hours ago [-]
Oh yeah. I forgot the geese are still at play here
largbae 3 hours ago [-]
Is this about Kimi or about Claude? Anthropic is circulating a $200B 2028 revenue target ahead of their IPO. If OpenAI plans to stay private longer, which their recent liquidity event might suggest, why not try and kneecap their competition?
claiir 10 hours ago [-]
Since it's only discounted on the standard "OpenAI," non-ZDR route (old pricing on Azure), I'm guessing a lot of users won't see this benefit? Since a lot of users enable a global "ZDR-only" toggle on OR
stavros 8 hours ago [-]
It would seem that getting lots of data is exactly the reason to discount this.
dannyw 6 hours ago [-]
OpenAI says they don't use any API data for training.
(there's probably going to be a reply about 'but how can you trust them'; I'm just stating what they say)
solenoid0937 6 hours ago [-]
Unless it's in your contract, they will use the data. They might not be using it now, but they will eventually.
stavros 6 hours ago [-]
Does OpenRouter say the same?
m4rtink 16 hours ago [-]
Price wars did wonders for many businesses, like the bike sharing industry in China.
Overgrown datacenters or mounds of GPUs dumped into the harbour next ?
Moto7451 16 hours ago [-]
I would in such a scenario expect the GPUs to be dumped to industrial breakers who would send them to China for refurbishment and repackaging before being sold again on Amazon, AliExpress, and Taobao as last gen gaming cards from weird brands and specs.
This is what happened after the great crypto GPU dumping.
dawnerd 13 hours ago [-]
Honestly can't wait for that to happen, same with memory, drives, etc. There's going to be a massive amount of server pulls hitting the market.
Joel_Mckay 10 hours ago [-]
The e-waste recyclers are pretty low on the pecking order, as the creditors will be first to strip these places for assets as Leopold Aschenbrenner discovered. =3
m4rtink 16 hours ago [-]
Yeah, I ment it as a joke - I agree with you. Watched the Gamers Nexus GPU investigation recently, where they were shown how a chinese soldering shop can transplant GPU chips to a new board, including memory chip reuse.
Hopefully we can look forward to all that useless datacenter AI crap gets repurposed in a similar manner into something actually useful for users.
kajaktum 16 hours ago [-]
I wouldnt be so hopeful because they dont use commodity hardware afaik
How are you going to use a h100 at home?
HeWhoLurksLate 14 hours ago [-]
with the same tricks gamers have always used? Modded drivers, undervolting, etc.?
m4rtink 7 hours ago [-]
Exactly + you can harvest the HBM memory on chip level & resolder on custom boards usable in PCs as VRAM (with scavenged GPU) or as normal RAM sticks.
skohan 11 hours ago [-]
I would happily buy up a load of datacenter GPU's at deep discount
Joel_Mckay 10 hours ago [-]
Won't have to wait very long... as they are eating their own already.
The Shrek movie market correction correlation may be due again in July 2027. =3
throwatdem12311 13 hours ago [-]
Mountains of GPUs next to the ET games in the landfill.
vatsachak 14 hours ago [-]
Well one person can use at most one bicycle at a time.
One person can use as many GPUs as they want.
throwatdem12311 13 hours ago [-]
At this point the models are “good enough” and whoever wins long term is gonna be whoever is the cheapest.
That’s why Chinese models are gaining traction and it’ll be the only way for OpenAI or Anthropic to keep up.
stillpointlab 10 hours ago [-]
I like to see this. I still prefer Fable (marginally) but my last big task was 100% Codex using Sol max (re-sizing my AWS infrastructure using CDK) and it did a very good job. No complaints, I could use this model happily to do what I need to get done.
If this nudges Anthropic to give me more Fable usage, that's even better.
0xbadcafebee 10 hours ago [-]
> my last big task was 100% Codex using Sol max (re-sizing my AWS infrastructure using CDK)
Fwiw, you could do this with any small or medium model, and it's easier with the aws-docs mcp. AWS is pretty stable, well documented, and programmatic, so most AI can figure out what it needs pretty quick
stillpointlab 40 minutes ago [-]
I'm not claiming an eval on what is or isn't possible. Rather I am stating my satisfaction with what I actually used and the actual result I obtained.
What I can say is that over the course of ~1 week I was able to review, plan, implement, test and release a significant change to a production system using Code Sol max. Any other claim about how any other model might have completed the same task is outside of my experience.
johnnyApplePRNG 5 hours ago [-]
OpenAI is making some really boneheaded moves these days.
They're reacting instead of leading, basically.
Cutting API prices 50% while millions of your paying subscribers have had their limits slashed and are all literally looking at the salivatingly-cheap chinese API prices availalbe on openrouter...
Not only did OpenAI and all of their cash somehow MISS the opportunity to purchase OpenRouter ...
Now they're giving a discount on an API that nobody even uses (get real, nobody's paying API prices to OpenAI ...
I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Needless to say, I will not be spending another dime with Codex or OpenAI.
This entire Codex reset limit fiasco has taught me they are not to be trusted.
Deepseek, here I come.
raincole 5 hours ago [-]
That's some bonehead take. For very starter, neither the other AI companies nor the customers will trust OpenRouter if it's owned by OpenA. It'd be squeezed to death from both sides. The only reasonable way for OpenAI's investors to have a share of OpenRouter is to invest directly, not via OpenAI.
> I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Yeah, sorry, skill issue. If you let AI run wild (if one session is $680 yeah it ran pretty wild) don't complain how it messed up your codebase.
enraged_camel 10 minutes ago [-]
>> Yeah, sorry, skill issue. If you let AI run wild (if one session is $680 yeah it ran pretty wild) don't complain how it messed up your codebase.
Not the OP, but I consider myself a skilled and heavy AI user, with multiple subscriptions in both platforms, plus OpenRouter.
A couple of weeks ago I gave 5.6 Sol a small/medium sized ticket to simplify part of the auth system. The ticket had a lot of details another Sol agent had collected during an exploratory session, and it was all vetted by Opus 5. I thought to myself that the implementation agent should have everything it needs. I still had it write a plan just in case, read the plan, made sure it matched the ticket, then clicked Approve and walked away.
I came back later that afternoon to a horror show. The agent had written 25,000+ LoC in the worktree. After 15 minutes of skimming through it, I realized that it had made the specced change, then convinced itself that it needed stronger verification, and over a series of compaction cycles ended up writing a static analysis harness so that it could prove that the change would be safe. Total bonkers.
Except, according to another Sol agent I showed the worktree to, the harness didn't actually do what the original agent claimed. The review agent said 98% of the worktree's code should be thrown away, and only the fix and its relevant unit and integration tests should be retained. I also asked Opus 5, and it theorized that 5.6 Sol must have gone through too many compaction cycles and lost track of its original goal.
This never happens to me with Claude models. Yes they write a lot of code and verbose comments, but I've never had a situation where a ticket that should take several hundred LoCs ended up with tens of thousands. When Claude overengineers something, I catch it during the planning phase, and it implements plans faithfully.
5.6 Sol is simply unreliable. It's too relentless and doesn't know when to stop. That's probably what caused the OP's $680 incident. I find it fascinating that people like it so much.
m101 4 hours ago [-]
I think what is going on here is that OpenAI are employing price differentiation to capture more of the market. The captive API customers are already there. There are bunch of potential customers that are price sensitive and are at openrouter. This move allows them to capture more of the market and increase profits (yes, they might be increasing profits at this level)
johnnyApplePRNG 3 hours ago [-]
You don't just jump from model to model on openrouter because somebody slashed prices, though.
People actually have to select and want to use Sol 5.6 in their routing.
dewey 4 minutes ago [-]
One of the selling points from openrouter is exactly that though, it makes it extremely easy to jump from model to model. For some background tasks for example I setup all the free models on openrouter and if one is not available it jumps to the next one.
m101 50 minutes ago [-]
Perhaps not most but surely some, and in the future perhaps increasingly easily.
I’d have thought that even today people would validate a number of models for certain tasks and on a daily basis go for the cheapest provider when they run that task.
krzyk 11 hours ago [-]
Is this pricing change only for openrouter? I don't see official OpenAI info about this.
Bombthecat 10 hours ago [-]
I was wondering the same, and it clearly says : 50% off, aka a sale, not normal price cut.
I don't get this thread.... Really. Is it full of bots?
dgunay 16 hours ago [-]
I'm loving this race to the bottom.
infinite_spin 15 hours ago [-]
I'm not having that experience. So far each major model update has been at least slightly better than the last, in ways I've found useful. Can't say it's perfect, or able to do exactly what I want without a decent amount of instruction/implementation/docs, but it's been useful enough to keep paying for it.
fn-mote 15 hours ago [-]
GP means race to the bottom in price not quality.
dgunay 13 hours ago [-]
Oh no the models are absolutely getting better, I'm just amazed that only 6 months ago I was using gpt-5.3-codex, and now I can use gpt-5.6-luna for similar results at like 1/15th the cost. Now 5.6-sol is being slashed by 50%? Amazing.
To me, it looks like the leading labs are investing more and more to improve their frontier models, and are being forced to charge less and less for them due to competitive pressure.
Maybe I'm wrong, but "reasoning as a service" is looking more and more like a... commodity.
egorfine 8 hours ago [-]
Slightly unrelated: what's up with the "tps" value? Does GPT-5.6 Sol really deliver just 32 tokens/second?
cmiles8 7 hours ago [-]
This is the opening salvos of an all out token price war.
With models a commodity at this point there isn’t much leverage for the big labs to keep their pricing anywhere near where it’s at. And that’s at the worst possible time as they need to be dramatically raising prices to have a viable business model.
Expect pricing to rapidly fall towards the underlying cost of compute and as players get really desperate we’ll likely see inference at less than the cost of compute as the market starts to rationalize and squeeze out weaker players who’s only play left will be to be the cheapest option in town.
The AI bubble is just waiting for the first player to scream mercy and cut capex as they simply can’t afford to throw more cash on the burning pile. That will be the trigger that implodes this bubble.
bwfan123 1 hours ago [-]
> The AI bubble is just waiting for the first player to scream mercy and cut capex as they simply can’t afford to throw more cash on the burning pile. That will be the trigger that implodes this bubble
Deepseek v4 and Kimi k3 have tightened the screws on the frontier models. With open-weights, anyone can host these for cost of compute. So, there is zero leverage left for the frontier labs.
What implodes it in my view is that enterprise adoption will stall. Enterprises are struggling to actually use these things in real workflows outside of coding and support.
fatata123 5 hours ago [-]
[dead]
matheusmoreira 11 hours ago [-]
Does this mean less subscription credit usage as well?
kaycey2022 9 hours ago [-]
No because i am still losing 50% of my weekly quota using sol on high.
Tadpole9181 10 hours ago [-]
It just looks like OpenRouter is doing a 50% sale on some models right now? Until OpenAI makes an announcement, I would assume no.
0xbadcafebee 10 hours ago [-]
OpenRouter doesn't do sales, they charge a premium, which is a flat 5.5% taken out of your credits. If you see something cheap on OpenRouter, it's because that one provider lowered its price. (Actually, correction, they will take 0.5% off their fee if you allow them to train your content)
Another thing some people don't notice is flex pricing, which is way lower than default pricing, for slightly worse latency and reliability. Depends on the provider and model
drivebyhooting 14 hours ago [-]
Has anyone had mixed experience running Ultra with and without /goal?
I come back to it after 8 hours to find it got stuck navel gazing imagined and Byzantine errors.
josh-wrale 17 hours ago [-]
Is this motivated by the value of the thinking traces gleaned from the traffic?
killingtime74 15 hours ago [-]
The thinking traces are server-side, not exposed
ec109685 15 hours ago [-]
They can’t decrypt the thinking traces.
dannyw 14 hours ago [-]
You can train a LLM to inverse summarised thinking into thinking text. It’s not perfect, but it gets you maybe 80% of the quality with proper techniques.
FWIW, there’s not that much value protected here anyway IMHO, and even raw thinking text can lie (as shown by Anthropic’s amazing research), so for legitimate interpretability research it’s limited.
Scaling frontier performance hasn’t been SFT-bounded for a while now; it’s now basically how much you can scale RL rollouts.
johnnyApplePRNG 5 hours ago [-]
The really frightening part for OpenAI, should be that apparently nobody seems to care. [0]
Their token usage on 5.6 Sol isn't even expected to double today.
I don't think you can expect a significant fraction of people to immediately switch based on a (presumably temporary) discount.
korijn 5 hours ago [-]
Perhaps people have grown tired of repeated bait and switch tactics?
ronfriedhaber 6 hours ago [-]
Hard to estimate what enabled the price cuts,
Yet OpenAI is doing some magic work, especially recently.
dannyw 6 hours ago [-]
Hard to estimate? Everyone knows the elephant in the room: capable open weight models.
wahid_seddiqi 1 hours ago [-]
This is honestly impressive. Cutting the price by 50% while pushing a model this capable is exactly the kind of move that makes advanced AI feel genuinely accessible.
Why though, to AB test/see the impact of a price cut on a platform with multiple competitors?
paxys 15 hours ago [-]
Is this captive audience not going to switch providers for a 50% discount? Especially when the effort is simply swapping one URL for another?
OutOfHere 15 hours ago [-]
OpenRouter is likely just leveraging Codex subscriptions.
prime_ursid 15 hours ago [-]
Wouldn’t that be against TOS?
OutOfHere 13 hours ago [-]
It might be through a level of indirection via an intermediate provider, offloading the TOS issue to the intermediate provider who couldn't care less.
matchagaucho 16 hours ago [-]
Right? Should we switch from direct OpenAI API integration to OpenRouter?
What's the incentive here?
Open Responses API doesn't appear to support state management (yet)
user43928 10 hours ago [-]
OpenRouter and the Vercel AI Gateway.
So yes, presumably a very small share of their total traffic.
baimoqilin 15 hours ago [-]
[dead]
meerita 4 hours ago [-]
This is great news for everyone. We should celebrate they're really doing price competition.
lyjackal 15 hours ago [-]
I saw this for Luna and then looked at the uptime and it said 85%. My interpretation is that this is just a gimmick where they serve the OpenAI flex tier at the same discount OpenAI provides for flex and then fall back to azure
tartakovsky 16 hours ago [-]
No ZDR. No dice.
bnrdr 6 hours ago [-]
dang, to avoid confusion from the title perhaps this should be edited to: “OpenRouter temporarily cutting GPT-5.6 Sol pricing by 50%”
dannyw 6 hours ago [-]
That suggests this is being funded by OpenRouter, and there's no indication this is the case (and I doubt OpenRouter can afford it; why would they anyway).
bnrdr 5 hours ago [-]
There also doesn’t seem to be an official indication that this is funded by OpenAI, hence the confusion in the thread.
therepanic 14 hours ago [-]
Even at these prices, switching from subsidized subscriptions to the API just isn't worth it. Not even close.
11 hours ago [-]
oliveralbertini 2 hours ago [-]
it's cut for short context
jeffybefffy519 13 hours ago [-]
Reading the comments in this thread, i honestly dont get it. 5.6-sol has felt like a regression in capability. In fact, every model since 5.3-codex has been a regression from OpenAI. I just find 5.6-Sol over engineers problems, takes absolutely ages to solve basic problems....
At this point, I'm considering going back to cursor over codex due to the ability to get more control over what model I use since there is clearly a heap of user preference and having frontier providers constantly shift the goal post with "State of the Art" is complete non-sense.
jeswin 12 hours ago [-]
It depends on what effort you're using etc. As an example [1] of what codex is capable of, here's hugo (written in golang) ported to TypeScript - and then a TypeScript to Rust transpiler which converts arbitrary TypeScript into Rust.
The TypeScript code which was transpiled into Rust (and is compatible with most hugo templates) runs faster than the original hugo.
The transpiler is still WIP, but the fact that it can do this says a lot of about how far LLMs have come.
jeffybefffy519 6 hours ago [-]
I find all effort levels of sol are the same in terms of amount of hallucinated unnecessary changes. Luna is much better all round on xhigh but my point still stands, every release of these new models is not an upgrade, its re-learning how to work with it.
Its like rehiring an employee every few months then training them up. Its honestly tiring and cant stay like this.
Opus has the same problem too…
jeswin 6 hours ago [-]
That's not been my experience. My prompting methods haven't changed much between recent GPT releases. I do put a lot of effort into building tooling and tests around a project, so the LLM output is converging around it.
They were built specifically for testing the C# target. There are several other large projects we built specifically for e2e testing.
But more interesting would be the tooling built to support this. For example, our current TypeScript parser [1] is a file-by-file port of Microsoft's TypeScript V7 compiler written in golang. The challenge here is that every time Microsoft changes code, we'll have to fix our code and tests. It's doable, but a fair amount of work.
So we decided to write tooling to transpile Microsoft's v7 compiler from golang, and autogenerate our compiler. That tool is called gotots [2] - and it already produces a fully working TypeScript compiler. It's 3x slower than TypeScript v6 compiler, but we hope to get to rough performance parity in a week or so. Everytime Microsoft makes an update, we run gotots and our parser gets updated as well.
My general point is that tests and tooling is tremendous value, and they are guardrails for LLMs to converge. I could have, for example, chosen not to write the go-to-ts transpiler, and live with porting Microsoft's parser line by line. But making such tools is something LLMs are good at, so it's a tradeoff well worth making. And the upside is that you don't have to use LLMs to port Microsoft's parser/compiler (a large and complex project) line by line.
tonyhart7 13 hours ago [-]
its over engineered problem solver ???? well because its a designed to do that
if you want to solve basic problem then use Luna
jeffybefffy519 7 hours ago [-]
I mean it added additional changes when it doesnt need to. Its basically hallucinating changes it thinks it needs to make regardless of effort levels i try.
SadErn 13 hours ago [-]
[dead]
hk__2 8 hours ago [-]
In my experience, "Sol" stands for "Stupid overengineering LLM". I’ve tried it at low/medium/high/xhigh effort levels and after a while I always end up to regretting my switch from Opus/Fable.
panda008 5 hours ago [-]
Sol is really stable for me.
bigbluedots 9 hours ago [-]
These threads seem to have become exceedingly vibes-based.
Yes, something may now be cheaper or more expensive or whatever, but there is no way to objectively measure quality (except for "trust me bro" benchmarks). So the discourse is people saying that for them, this or that model was better - which is a very low value data point.
pessimizer 33 minutes ago [-]
It's got to be because AI has become a commodity, people are looking to get the best value for money, and things change based on your specific application (which may be nearly unique) and literally the time of day.
It's a bunch of bread bakers talking about wheat suppliers.
oblio 9 hours ago [-]
> These threads seem to have become exceedingly vibes-based.
If you remember programming language discussions, they are exactly like this.
Software development is still in the leeches and bloodlettings phase.
bigbluedots 8 hours ago [-]
Yes, programming language discussions can be vibes-based and therefore low value too, but sometimes the more concrete aspects of the languages at hand, e.g. language features and tradeoffs are discussed. That is something that I'm not seeing in equivalent AI discussions.. there is a lot of how a particular AI model "feels" to interact with.
ComputerGuru 15 hours ago [-]
Does OpenRouter eat this cost to get their hands on a copy of the conversations people are using with the model?
9cb14c1ec0 15 hours ago [-]
No, this is OpenAI doing the discount, not Openrouter by themselves. OpenAI is crushing it with their 5.6 models, and they probably decided there was no better time to grab as much market share as possible.
ComputerGuru 4 hours ago [-]
But the discount is only available via OpenRouter.
Noaidi 14 hours ago [-]
I don’t understand this at all. They have never been profitable yet. How is this helping them? When it be more likely the case that not enough, people are using it as the prices they established already? So now they have to lower the prices?
ipaddr 12 hours ago [-]
You lower prices for marketshare. Fable became a mythological model to leadership because they were the first story of ai escaping and hacking another company. The it's so dangerous the public can't use it narrative is sticky so OpenAI is showing off its model so as many eyeballs as possible. We're in the samples in the supermarket phase.
nprateem 12 hours ago [-]
They're prepping to IPO. They want top level metrics like usage they can use to pump investors, not nonsense like profitability.
voiper1 11 hours ago [-]
OpenRouter offers 1% discount to save your conversations, explicitly opt-in.
Anything else they don't save it. Even if they tell you the model provider saves your data for training.
ee334y5rthsrth 5 hours ago [-]
This could cause the AI market to crash. If companies can't make money off their users, their entire financial plan will collapse and their stock prices will plummet.
asd000hh 3 hours ago [-]
Coool
dvrp 15 hours ago [-]
For context, Stripe has just acquired OpenRouter for >$7B.
I’d bet that explains this move!
indigodaddy 15 hours ago [-]
Why would the potential acquisition have anything to do with this? They do discounts all the time on various models. Luna was 50% off last week..
gip 15 hours ago [-]
Not sure as OpenAI models (Sol, Luna,..) are also discounted on the Vercel AI Gateway rn. My bet is on OpenAI trying to drive more enterprise customers to their models through API.
Topology1 13 hours ago [-]
How can they do this? Are they subsidizing it out of pocket?
ben8bit 11 hours ago [-]
Terra is also a fantastic model.
timedude 3 hours ago [-]
When are RAM prices getting cuts tho. It is getting fucking ridiculous.
shevy-java 11 hours ago [-]
They are really getting desperate. The bubble is coming closer to an end here.
gutterscale 10 hours ago [-]
GPT-5.6 sol starting to be a real workhorse at this price point
vorpalhex 17 hours ago [-]
Do other people find 5.6 to be worse at most simple tasks and frequently over complicate things?
I asked it to write a user todo and it turned out a four page essay. I gave the same task to 5.4 and got the small list of checkboxes I expected.
qup 16 hours ago [-]
I've found it to be great for planning code changes (or new projects). I use the superpowers plug-in which I think guides the planning.
Then I switch models (to luna) before implementation. I find this combo nearly always does what I want.
I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above.
I like that skills are easy to try out, too.
jeremyjh 13 hours ago [-]
I stopped using superpowers because it wanted to turn every tiny bug fix into a $37MM DOD project. I got effective results but it took ages. I may try again - I need to find a good way to run different profiles in my harness so I can easily shut it off. The default planning workflow in OMP is pretty good though.
I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if the task is well defined and straightforward, but there are lots of models now that you can say that about.
jcastro 14 hours ago [-]
I have the same setup you have, love it!
drdexebtjl 15 hours ago [-]
You would probably get better results with Luna for the real simple tasks, or Sol with low thinking effort.
I find that I get exactly the effort that I asked for, which is pretty nice. The other side of that coin is that these are the least lazy models I’ve used so far. They will go on elaborate tangents to complete the task when I want them to.
khacvy 13 hours ago [-]
any good best practice for effort selection on claude. I always use high as default.
drdexebtjl 3 hours ago [-]
Opus 4.8 on Low is similar in price and quality to Sonnet 5 on High, but much faster.
But otherwise I don’t use Claude anymore.
infinite_spin 15 hours ago [-]
I've found it's worse for simple tasks too, and I have to give it stricter guidelines, and sometimes it doesn't follow the same patterns I've grown to expect. I've found using 5.6 (sol) is good for diagnosing issues though, especially in terms of optimization of some given path
dimgl 16 hours ago [-]
Yep. I have not yet had a single good experience with Sol or the 5.6 models on a variety of harnesses and configurations. It overthinks, overcomplicates and often makes my code into an unmaintainable sludge. It'll usually take 5+ turns of steering to get it in the right direction.
OutOfHere 17 hours ago [-]
It's your responsibility to set an appropriate level of Thinking. For simple tasks, I use the instant model. As an approximation, the choice is proportional to the amount of time I want it spending on the task. Also, you can always ask it to respond succinctly.
code_biologist 17 hours ago [-]
[dead]
kristo 8 hours ago [-]
It shocks me how little people seem to care that they are supporting an evil Zionist lizard man who molested his sister and is happy supporting trump. Doesn’t even come up in the conversation here. I don’t really care if sol is a bit better, I still make decisions on more than that.
Is the HN community just too online and sucked in to the musk mind manipulation vortex? Or what is going on? Why does nobody seem to care?
aetherspawn 10 hours ago [-]
Can we get it for the reduced rate direct from OpenAI though?
gxs 13 hours ago [-]
Absolutely not
I’ve used Claude exclusively for the past few months
Was excited when Sol came out a few weeks ago and loaded it up
I made the mistake of treating it as if it were Claude - I’d assumed they were close enough in ability and treated them that way
Well, turns out my instruction sets for Claude are 100% too complicated for Sol
Sol made the stupidest assumptions, constantly did things that it wasn’t asked to do and always approached code in what I considered a weird way - I had redo a lot of my prompts to get it anywhere close
Now, did it do good work?
Yes, on occasion. But with LLMs and coding, consistency is the name of the game. Constantly having to correct the LLM and constantly feeling paranoid that it won’t listen makes for an exhausting session
Maybe if you “came up” in the codex world you’re more fluent with it, but sticking with Claude for now
iammrpayments 12 hours ago [-]
Why are you being downvoted, is this post an ad or something.
cbg0 11 hours ago [-]
Probably because it's PEBCAK.
gxs 10 hours ago [-]
Hard to take offense from someone who uses pebkac seriously
Kudos to you though for being your authentic self so publicly
dana321 8 hours ago [-]
"rewrite unreal engine in rust, make no mistakes"
gxs 34 minutes ago [-]
That’s exactly what I did! How’d you know?
Didn’t mean to make you upset sorry
chinagenie_ai 7 hours ago [-]
[flagged]
floki165 9 hours ago [-]
[flagged]
mohammedmsgm 8 hours ago [-]
[dead]
xcupapps 4 hours ago [-]
[dead]
Bob_bo 11 hours ago [-]
[dead]
senectus1 15 hours ago [-]
[dead]
Taikhoom10 14 hours ago [-]
[flagged]
Scene_Cast2 14 hours ago [-]
Oh hey, that's cheaper than Kimi K3! Amusing to see a SOTA OpenAI model be cheaper than a Chinese open weight model.
Fwiw I love K3 and use it as a daily driver. I haven't tried Sol, as I dislike OpenAI.
ardel95 8 hours ago [-]
My bet is that OpenRouter began steering GPT-5.6-sol users towards flex tier, which is already 50% off.
So this isn’t really a price cut. As to why, lots of possible reasons. Perhaps an agreement with OpenAI to help them drive up more diverse traffic priorities.
Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.
(Makes sense to me, just curious about this additional piece.)
Looking at reserved capacity cost for PTUs on azure, which I think they’d probably not subsidize but can’t be sure, I’m inclined to not agree with the vast undercharging for tokens hypothesis.
https://zenmux.ai/deepseek/deepseek-v4-flash
[1] https://youtu.be/kacf2bib-X0
Well, DeepSeek just raised prices.
I assume it's highly use case dependent, though?
Even before the price cut seems like Sol was price competitive with Kimi
https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhig...
And now it should be considerably cheaper
You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.
If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.
It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents.
If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive.
And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.
Surely you don't want them to be the reason the bubble bursts?
Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).
However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.
I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.
There's people that have tried to contact Jeff McBride and follow the IP trail but the IP is currently owned by a company that went defunct. Not sold, but no one is even bothering to register its LLC any more, it's simply dead.
Some of the topics it’s flagged have been hard for me to understand what it seeing that can be remotely concerning in my requests.
Sol is great and has never blocked a request, and generally gives great answers. Happily switched over to it now.
Which is definitely protecting their turf, but also probably a little bit hiding their “RSI” abilities for competitive reasons. My theory is that a lot of “safety blocking” is actually WIP training of new business directions. Anthropic has started hiring biologists and has opened a preview of a “Claude code for bioinformatics”. I’m guessing they’re tweaking their bioinformatics market play, and block “bio safety” requests so competitors can’t learn about their training.
That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.
Compared to what I am doing at home experimentally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.
But I wouldn't trust lower tier models for end to end solutions.
Getting the AI to output code that you like is difficult.
As an example, let's say in React you have a "useLocale()" hook.
The AI will happily pass down locale as a prop to 5 child components instead of just calling the hook in the component.
A review from another model did not flag such stylistic issues either.
I believe that the latest models are very good at functionally achieving the goal, but still have poor taste for UX or code quality.
The most productive use of AI for software development happens in an environment where you do not review the code but test the UX end to end.
I think it sometimes worked, for example for testing preferences, but sometimes it did not.
Could be a problem with the harness also.
In any case, I feel that it's a bit playing whac-a-mole with explicit rules for things that a more intelligent model should do by default.
MAI also offers a ultra cheap version that's competitive with Luna.
So much so that the models look like they were designed by a product manager explicitly to eat away OpenAI's market share.
Vscode even pushed them quite hard onto users with the latest release, going to the extent of putting up a modal to convince users to try them out.
The only obvious objection I can think of to this line of thinking (at least from a technical perspective) is "how does someone build up the knowledge to be able to use a tool effectively in that way if not by doing things by hand at first?" The honest answer that is "I don't know, but that's also pretty much exactly the type of thing my employers have never been paying me to solve in the first place". Even just a decade into my career, there have already been plenty of times in my career I've struggle to convince people that we should do stuff in a way that won't bite us in the ass a month or two down the line, and in the times I've managed to succeed, it's usually only by putting in more of my own time and effort to make the initial investment seem more palatable. Luckily right now I'm not in one of those times when I'm having to go full throttle to keep the lights on a few months from now, but I don't have enough fuel in reserves to work on a plan for when we need to build a new rocket in another ten years. Maybe ask me next month.
It's just vibes
But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.
When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ, academic tests). This is not easy.
It's... really just vibes?
Always has been.
After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.
Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value by far and there's no contest.
---
Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.
Both models suffer from doing-too-much. But both models are fundamentally really smart and knowledgeable. I think it's really close and pricing cuts really spice things up for us consumers! Sol is a clear winner in the value department and the $100 plan is enticing!
---
*Claude Code usage is reducing by 33% in 2 days, Wednesday August 19... cmon anthropic: clau.de/cc-50-promo
However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).
One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.
Given the 50% discount on Sol and how smart it is, yeah it's unprecedented value. If you only want to use low effort, there's a clear winner here on value and it's not even close!
Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?
For small tasks, you can just use something like low or medium effort and it can usually avoid mistakes; after all, the model will test the code anyways and can do some baseline level of iterating.
In regards to cost, we need to acknowledge how generous OpenAI was in the last couple months with Codex usage credits (no weekly limits) and usage resets. It afforded me many a dive with Codex! Yes it uses more tokens, but sometimes it's worth it -- just depends on what you're working on.
Finally, Ultra(code) isn't that bad when it comes to cached tokens. I think folks overstate the general token usage of ultra effort on both providers.
---
Both models are great at green-fielding a project when given detailed specs.
Both models overthink too liberally (imo) during these larger multi-shots. Sol overthinks more than Fable.
Both models are really smart and perform great for general knowledge and regular coding tasks.
I'm in Australia, and Fable downgrades to Opus when testing for bugs in memory in a legacy C code base. If Fable starts taking initiative and writes a test case that involves writing to a null pointer, that's the end of the conversation.
> even after completing the verification program
Was it easy to complete it?
I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.
Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.
Also Opus 5 is fine if your codebase is simple.
I hear you on the downgrades, I'm 13/13 on downgrades, and last downgraded me to Sonnet for asking for reasoning chain.
I just can't get Opus (Opus 5 is dumb as a rock, to be fair) or Sol to do that, so I exclusively use Fable for personal work. When I reach my weekly limit, usually on the last day close to the reset, I just go back to coding by hand ¯\_(ツ)_/¯
Heck, it even does security reviews and fixes, as long as I don't ask it to "attack" the codebase. I'm planning on using Kimi or GLM for that part.
It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.
I don't even want to see the codebases this model is fucking up.
It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.
What on earth are you asking it?
I have witnessed 5.6 Sol Ultra edit line after line of literally empty lines ... for hours.
I wasn't literally watching it, I came back to a goal (that it started for itself without my approval!) that had done nothing but that for some reason.
It couldn't explain why it had started.
I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.
Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.
cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).
what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...
for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.
Its a crutch that is no longer competitive
They STILL don't have an option to "Sign in with Apple" on the website, but they do for Google??!? (and on iPhone of course)
Screw that asinine UX
(and no it wasn't better than Codex at this particular task)
Can somebody at Anthropic tag claude in slack or whatever goofy shit you do and ask it to add Apple OAuth to your website? Clearly humans aren't testing it.
and sure enough, I was right to do so: They don't even let you remove your payment method afterwards. Every other store, Steam etc., lets you.
No way I have enough trust to install their desktop app after that, so I just want to try it through their website..
Can Sign In with Google, but not with Apple
so you gotta open the Passwords app, copy your random email, paste into the website, then copy the OTP from your email..
It's been that way for at least a year
The desktop app was clunky too the last couple times I tried it a few months ago
and the AI itself hasn't been that hot compared to ChatGPT/Codex either: https://i.imgur.com/jYawPDY.png
So all the Claude hype posted on HN seems like a case of the emperor with no clothes to me
(P.S. The thing I just now tried to do on Claude hit the weekly usage limit after 2 minutes)
If Sol isn't the best model, it is up there...
You don't cut the price of the best model for no reason...
Always has been. My prediction is that both OpenAI and Claude will go bust unless they deliver a killer product. And unlike scrappy startups, they have a pretty serious deadline because creditors will come a-knockin'.
There's little to no functional difference between Kimi, Qwen, Sol, Opus, etc. All flagship models are within like 1-5% of each other and the real moat will be what's always been the hard part: making a good product.
https://epoch.ai/data/ai-companies?view=graph&tab=revenue
The Chinese models are cheap because no one is using them. But they can't actually afford (or have capacity) to serve enough people to kill the giants. This is evidenced by them all recently hiking prices or limiting usage.
It's possible that they build out in China at an unreal pace, China doesn't have concept of "community input" to drag down state projects, but then you are left giving your IP to China. Just ask western hardware businesses how well that goes.
If compute is the moat, that doesn’t really make their position any less precarious.
They're guaranteed to get over whatever hump you think they're in unironically. Uber/Tesla have been in far worse situations and despite Elon being an idiot/liar you see how they performed when even the most bullish of investors called for their heads
Don't know about that.
I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc.
Ran a complete Fable/max code review. Took over a month on a subscription. Now I've switched to OpenAI and am repeating the exact same review with Sol/max.
It's still not done yet but preliminary findings suggest Sol can only reproduce 70-90% of Fable's findings. So I think these models aren't as close as we've been led to believe.
Claude already has a killer product (claude.ai/chat is a Swiss army knife) but just relying on people typing stuff into chat is not enough to sustain the company.
The other strategy is entrenching yourself as the LLM of choice into existing products (like ChatGPT is on Apple products).
Depends on your use case. the Chinese models are not there yet.
This is basically undercutting KimiK3 and Grok 4.6 where previously utilised gad soke advantages but was a step more expensive
https://trackingai.org/home
That shows a bunch of models, including Sol, with a discount. None of them say how long it’s for, but I’d assume in all their cases it’s for a limited time as the banner said, and only on OpenRouter.
I posted elsewhere, but the Azure uptime & performance for Sol is truly dire. OpenAI is offering 5x faster latency, 4x faster tokens generation, and vastly better uptime (Azure US has only 87% uptime), all for 50% of the price now. I assume the pricing is to compete with other shiny new models (Grok, Qwen etc), but it might also be to cut-off a truly poorly performing Microsoft hosting experience.
Maybe they want to see how much market they can grab with Sol?
This might help but there are already cheaper models with Sol's intelligence more or less, the most notable being Grok 4.6 at $6/m which makes it a tougher sell
[0] https://tinfoil.sh
It's now my open weight inference provider of choice, since on top of the privacy/security characteristics it's also reasonably cheap.
tinfoil asks more than 10x the output cost ... $1.90 per 1M tokens instead of $0.18 per 1M tokens for my favorite model (Deepseek V4 Flash 0731) on my favorite provider (DeepInfra) currently, for example.
I am quite happy with them so far.
Batch API is a bit harder, not many models/providers support Batch. It primarily is only the Gemini/ChatGPT/Claude models that do. DigitalOcean does support a 50% discount on Batch API via them directly, not listed on OpenRouter.
[1] https://docs.digitalocean.com/products/inference/how-to/use-...
For Mythos and even Fable they require prompt retention on their end.
edit: or more precisely if you want to access Mythos/Fable ZDR does not apply, and depending on config the exclusion can affect other models.
But if their employer is bringing millions of dollars of potential spend to the table, can tell you from years of experience that turns a lot of 'no's' to 'yes'.
As for ZDR and court-orders, what would you rather happen there? Violate the law or comply with holding the data? I would bet that any ZDR agreement has this court-ordered risk mutually understood and agreed upon.
I'd be surprised if they closed out over 100B tokens today (they did 101B on sol 5.6 yesterday).
```
This changes how Claude Code communicates with you
```You can set up a custom one under `vi ~/.claude/output-styles/eli5.md` and `eli5` would show up in the list above:
```
--- name: ELI5 description: keep it simple please keep-coding-instructions: true ---
I can't process lengthy text, talk to me like I'm 5.
Small words, short sentences, short paragraphs. If you have to use a big word, explain it right after. Only return what's actually necessary. Just tell me what you did, did it work, what do I do now.
If I have to decide something: show 3 options max, the context I need to pick fast, and which one you'd go with.
Keep paths and commands exact. I have no brain cells left for the rest.
```
However, the fact that you can do this says something in my opinion. And it doesn't necessarily point to somewhere good if I'm being honest. These niche config options are kinda crazy? I've been using CC for the past year and just learned about this last week, I don't know if it's new, or whether it's been there a while, but I feel like it shouldn't be necessary.
God knows I've tried. I've got a variant of the ASD-STE100 trick which does the job, mostly, at the start … but get to about 100k of context and it goes out the window.
The model's personality is too strong for simple suggestion, alas.
Codex has always beaten claude in coding benchmarks, hasn't it?
Their api pricing is absurdly expensive.
I assume at this point that it subsidizes subscriptions.
I've gotten more work done on a second chatgpt pro $100/mo subscription than I did with ~$150 of paying for usage through the app.
Privacy, experimenting with ML and "unorthodox" needs are currently the only acceptable reasons to do local.
Experimentation and privacy are definitely advantages, but it's also quite a lot of fun.
Also if you don’t specify, most end up being the same as the parent model which is pretty wasteful.
I engineered a skill that spins up Terra High agents for most sub-agents, resorting to Sol Medium for technical research and Luna High for code/in-project research tasks.
On a slightly different topic, Luna Max is incredibly capable and doesn’t use as much quota (Luna tokens are dirt cheap).
No one can judge the enjoyment, learning, and hobby aspects. Just wondering if there is an end goal for that much overall expenditure (time, money, energy, etc.)
I think on average AI energy usage is not as big a deal as everyone is panicking about, but your usage is truly absurd and I don't know how you can live with that. It's immoral.
As for co2, it depends on the provider, it could be way lower as well.
As for ethics, you don’t know what he works on, and how effectively - he might be saving 10x that much of co2 for the planet.
I’m guessing that Wh/token estimate is several orders of magnitude too high.
Leaked financial documents from 2025 show the company reported an operating loss of approximately $20.9 billion against $13.1 billion in revenue.
Is any amount of tokenmaxxing moral?
Do some people still deny you can do a shit ton of work with AI?
I ended up borrowing my gf's phone number just so I could get access for work. Ridiculous
OpenAI's docs still show non-discounted pricing https://developers.openai.com/api/docs/models/gpt-5.6-sol
https://vercel.com/changelog/gpt-5-6-sol-is-50-off-on-ai-gat...
You're literally encouraging someone else to come in and steal your customer base,
Multiple times a year, retailers here in Australia have co-ordinated sales on Apple products. Apple.com or their retail stores don't have these sales.
But they're clearly Apple-funded when competing retailers launch the same sales on the same days; and the margins aren't enough for retailers to take a loss.
It also seems to be providing a vastly better user experience - Azure has less than 99% uptime (Azure USA only has 87% uptime), latency of 20 - 30 seconds, and a mere 8 tokens per second. OpenAI is offering 32 tokens per second (4x faster), 4 seconds latency (5x faster), and all for half the price of what Microsoft is charging for a vastly inferior experience.
Data taken from this page:
https://openrouter.ai/openai/gpt-5.6-sol#providers
Has OpenAI struck a deal with openrouter and that's why we're seeing preferred pricing?
Is openrouter taking a loss on sol API calls to grow adoption?
How temporary is the reduction in price?
Do you include research and training costs? Of all models or only the ones being served? What percent of the R&D budget do you allocate to inference? What about data center capacity? Do you count future commitments? All the circular financing deals? Do you count employee equity grants as costs? At what valuation?
We also have another solution for "whatever accounting decides": generally accepted accounting practices. It's far from perfect, but GAAP figures are what you should be looking at; not "adjusted GAAP" or whatever invention.
(there's probably going to be a reply about 'but how can you trust them'; I'm just stating what they say)
Overgrown datacenters or mounds of GPUs dumped into the harbour next ?
This is what happened after the great crypto GPU dumping.
Hopefully we can look forward to all that useless datacenter AI crap gets repurposed in a similar manner into something actually useful for users.
https://www.youtube.com/watch?v=rE75WvOtcu8
The Shrek movie market correction correlation may be due again in July 2027. =3
One person can use as many GPUs as they want.
That’s why Chinese models are gaining traction and it’ll be the only way for OpenAI or Anthropic to keep up.
If this nudges Anthropic to give me more Fable usage, that's even better.
Fwiw, you could do this with any small or medium model, and it's easier with the aws-docs mcp. AWS is pretty stable, well documented, and programmatic, so most AI can figure out what it needs pretty quick
What I can say is that over the course of ~1 week I was able to review, plan, implement, test and release a significant change to a production system using Code Sol max. Any other claim about how any other model might have completed the same task is outside of my experience.
They're reacting instead of leading, basically.
Cutting API prices 50% while millions of your paying subscribers have had their limits slashed and are all literally looking at the salivatingly-cheap chinese API prices availalbe on openrouter...
Not only did OpenAI and all of their cash somehow MISS the opportunity to purchase OpenRouter ...
Now they're giving a discount on an API that nobody even uses (get real, nobody's paying API prices to OpenAI ...
I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Needless to say, I will not be spending another dime with Codex or OpenAI.
This entire Codex reset limit fiasco has taught me they are not to be trusted.
Deepseek, here I come.
> I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Yeah, sorry, skill issue. If you let AI run wild (if one session is $680 yeah it ran pretty wild) don't complain how it messed up your codebase.
Not the OP, but I consider myself a skilled and heavy AI user, with multiple subscriptions in both platforms, plus OpenRouter.
A couple of weeks ago I gave 5.6 Sol a small/medium sized ticket to simplify part of the auth system. The ticket had a lot of details another Sol agent had collected during an exploratory session, and it was all vetted by Opus 5. I thought to myself that the implementation agent should have everything it needs. I still had it write a plan just in case, read the plan, made sure it matched the ticket, then clicked Approve and walked away.
I came back later that afternoon to a horror show. The agent had written 25,000+ LoC in the worktree. After 15 minutes of skimming through it, I realized that it had made the specced change, then convinced itself that it needed stronger verification, and over a series of compaction cycles ended up writing a static analysis harness so that it could prove that the change would be safe. Total bonkers.
Except, according to another Sol agent I showed the worktree to, the harness didn't actually do what the original agent claimed. The review agent said 98% of the worktree's code should be thrown away, and only the fix and its relevant unit and integration tests should be retained. I also asked Opus 5, and it theorized that 5.6 Sol must have gone through too many compaction cycles and lost track of its original goal.
This never happens to me with Claude models. Yes they write a lot of code and verbose comments, but I've never had a situation where a ticket that should take several hundred LoCs ended up with tens of thousands. When Claude overengineers something, I catch it during the planning phase, and it implements plans faithfully.
5.6 Sol is simply unreliable. It's too relentless and doesn't know when to stop. That's probably what caused the OP's $680 incident. I find it fascinating that people like it so much.
People actually have to select and want to use Sol 5.6 in their routing.
I’d have thought that even today people would validate a number of models for certain tasks and on a daily basis go for the cheapest provider when they run that task.
I don't get this thread.... Really. Is it full of bots?
Maybe I'm wrong, but "reasoning as a service" is looking more and more like a... commodity.
With models a commodity at this point there isn’t much leverage for the big labs to keep their pricing anywhere near where it’s at. And that’s at the worst possible time as they need to be dramatically raising prices to have a viable business model.
Expect pricing to rapidly fall towards the underlying cost of compute and as players get really desperate we’ll likely see inference at less than the cost of compute as the market starts to rationalize and squeeze out weaker players who’s only play left will be to be the cheapest option in town.
The AI bubble is just waiting for the first player to scream mercy and cut capex as they simply can’t afford to throw more cash on the burning pile. That will be the trigger that implodes this bubble.
Deepseek v4 and Kimi k3 have tightened the screws on the frontier models. With open-weights, anyone can host these for cost of compute. So, there is zero leverage left for the frontier labs.
What implodes it in my view is that enterprise adoption will stall. Enterprises are struggling to actually use these things in real workflows outside of coding and support.
Here are all the providers giving discounts: https://openrouter.ai/collections/discounted-models
Another thing some people don't notice is flex pricing, which is way lower than default pricing, for slightly worse latency and reliability. Depends on the provider and model
Paper: https://arxiv.org/abs/2603.07267
FWIW, there’s not that much value protected here anyway IMHO, and even raw thinking text can lie (as shown by Anthropic’s amazing research), so for legitimate interpretability research it’s limited.
Scaling frontier performance hasn’t been SFT-bounded for a while now; it’s now basically how much you can scale RL rollouts.
Their token usage on 5.6 Sol isn't even expected to double today.
[0] https://openrouter.ai/openai
OpenRouter attributes this promotion to OpenAI https://x.com/OpenRouter/status/2089416739398254662
What's the incentive here?
Open Responses API doesn't appear to support state management (yet)
So yes, presumably a very small share of their total traffic.
At this point, I'm considering going back to cursor over codex due to the ability to get more control over what model I use since there is clearly a heap of user preference and having frontier providers constantly shift the goal post with "State of the Art" is complete non-sense.
The TypeScript code which was transpiled into Rust (and is compatible with most hugo templates) runs faster than the original hugo.
[1]: https://github.com/tsoniclang/tsonic-examples/tree/main/rust...
The transpiler is still WIP, but the fact that it can do this says a lot of about how far LLMs have come.
Its like rehiring an employee every few months then training them up. Its honestly tiring and cant stay like this.
Opus has the same problem too…
This is the csharp target for tsonic (a TypeScript to C#/Rust/Python/Triton transpiler). It has a bunch of tests here: https://github.com/tsoniclang/tsonic-csharp/tree/main/test
More comprehensive e2e proving grounds are at
1: https://github.com/tsoniclang/proof-is-in-the-pudding
2: https://github.com/tsoniclang/tsumo/
They were built specifically for testing the C# target. There are several other large projects we built specifically for e2e testing.
But more interesting would be the tooling built to support this. For example, our current TypeScript parser [1] is a file-by-file port of Microsoft's TypeScript V7 compiler written in golang. The challenge here is that every time Microsoft changes code, we'll have to fix our code and tests. It's doable, but a fair amount of work.
So we decided to write tooling to transpile Microsoft's v7 compiler from golang, and autogenerate our compiler. That tool is called gotots [2] - and it already produces a fully working TypeScript compiler. It's 3x slower than TypeScript v6 compiler, but we hope to get to rough performance parity in a week or so. Everytime Microsoft makes an update, we run gotots and our parser gets updated as well.
[1]: The old parser - https://github.com/tsoniclang/tsts-legacy
[2]: Golang to TypeScript transpiler - https://github.com/tsoniclang/gotots
My general point is that tests and tooling is tremendous value, and they are guardrails for LLMs to converge. I could have, for example, chosen not to write the go-to-ts transpiler, and live with porting Microsoft's parser line by line. But making such tools is something LLMs are good at, so it's a tradeoff well worth making. And the upside is that you don't have to use LLMs to port Microsoft's parser/compiler (a large and complex project) line by line.
if you want to solve basic problem then use Luna
It's a bunch of bread bakers talking about wheat suppliers.
If you remember programming language discussions, they are exactly like this.
Software development is still in the leeches and bloodlettings phase.
Anything else they don't save it. Even if they tell you the model provider saves your data for training.
I’d bet that explains this move!
I asked it to write a user todo and it turned out a four page essay. I gave the same task to 5.4 and got the small list of checkboxes I expected.
Then I switch models (to luna) before implementation. I find this combo nearly always does what I want.
I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above.
I like that skills are easy to try out, too.
I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if the task is well defined and straightforward, but there are lots of models now that you can say that about.
I find that I get exactly the effort that I asked for, which is pretty nice. The other side of that coin is that these are the least lazy models I’ve used so far. They will go on elaborate tangents to complete the task when I want them to.
But otherwise I don’t use Claude anymore.
Is the HN community just too online and sucked in to the musk mind manipulation vortex? Or what is going on? Why does nobody seem to care?
I’ve used Claude exclusively for the past few months
Was excited when Sol came out a few weeks ago and loaded it up
I made the mistake of treating it as if it were Claude - I’d assumed they were close enough in ability and treated them that way
Well, turns out my instruction sets for Claude are 100% too complicated for Sol
Sol made the stupidest assumptions, constantly did things that it wasn’t asked to do and always approached code in what I considered a weird way - I had redo a lot of my prompts to get it anywhere close
Now, did it do good work?
Yes, on occasion. But with LLMs and coding, consistency is the name of the game. Constantly having to correct the LLM and constantly feeling paranoid that it won’t listen makes for an exhausting session
Maybe if you “came up” in the codex world you’re more fluent with it, but sticking with Claude for now
Kudos to you though for being your authentic self so publicly
Didn’t mean to make you upset sorry
Fwiw I love K3 and use it as a daily driver. I haven't tried Sol, as I dislike OpenAI.
So this isn’t really a price cut. As to why, lots of possible reasons. Perhaps an agreement with OpenAI to help them drive up more diverse traffic priorities.