

It’s also unavoidable.
Developer and refugee from Reddit


It’s also unavoidable.


(Non)Human centipede.


I’m waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.


“Hi everyone, welcome to my talk on fainting from anxiety, as triggered by… this… this… uuuuuuuhhhgg…” [Thud]
I have an autographed original drawing. Unfortunately, it’s a drawing of the Sandman, by Neil Gaiman. Yeah… I don’t have that on display in my house anymore.


If mods disagree that this is oniony enough, that’s fine with me too.
Y’all good here. This one really does read like satire.


I mean, did anyone expect anything different? The frontier providers that Copilot depends on switched to token-based billing in a desperate and ludicrous attempt to somehow turn a profit on AI, so of course Microsoft is gonna jack up their prices too.


There was a lot of news about the Muttsee Dam solar project a few years ago, so it’s a real thing and a good thing. But it’s hardly current news. It’s been operational since 2022.


It wasn’t. But that doesn’t mean we shouldn’t strive to make it be the best version of itself it could be, and fight tooth and nail against attempts to make it worse.


Very serious. Your personal amount of usage means nothing at all in this conversation. It is entirely about tokens per watt. The amount of energy the memory operations involve scale incredibly well when people are accessing the same object in memory simultaneously. Last I looked it was around a 10x difference for the same models efficiency.
Hold up. Are you talking about caching? Because if you are… yeah. That has nothing to do with the model and everything to do with the service layer around the model. The same service layers can be - and have been - implemented in tools like Lemonade Server, llama.cpp, Ollama, etc.
And I really do want to know your sources.
Mine say GPT 5.5 is probably using quite a lot more than 0.34 Wh per query (0.34 Wh is what Sam Altman claimed for the then-current version of GPT in June of 2025, but he hasn’t released numbers since then and no one has done an independent analysis). With Claude, an independent estimate from last year pegged Sonnet at 0.8 Wh for a short prompt, 2.8 Wh for a medium one, and 5.5 Wh for a long one. Current numbers are, again, almost certainly much higher. And just for fun, there’s DeepSeek (which I’ve never used and never would use), with the reasoning-tuned DeepSeek-R1 hitting a whopping 29 Wh for a complex query.
Meanwhile, small, open models are probably in the 0.07 - 0.2 range, depending on the model, the hardware it’s running on, and the nature of the query. Of course, there are much weightier open models too, with ones like Llama 3.1 405B using about 9 Wh for a medium-length prompt. On the other hand… who is going to run that on their local machine?
Look… If I’m wrong, and using local models the way I do - sparingly and infrequently - really does consume more electricity than using Claude Code, I want to know. I have no problem whatsoever with eschewing AI models entirely, since I despise all of them. But given how tight-lipped OpenAI and Anthropic are about energy consumption per average prompt, and what independent analyses have estimated, I am highly skeptical that they are acting as some sort of paragons of environmental stewardship.


You’re probably burning more energy turning it off and on again. It doesn’t really use any noticeable power sitting idle.
I am absolutely not burning more energy than a frontier model by doing things like putting my laptop to sleep or shutting down unused services when I want to conserve battery power.
Anyway, a direct comparison would be pretty difficult because your model is probably tens of billions of parameters, not over a trillion.
True.
Energy consumption per output token will probably be a bit higher for the frontier models but something that people have found is that higher quality models often need fewer tokens to achieve the same goal.
That’s actually not true. In fact it’s much the opposite. Frontier models churn through tokens at a much higher rate, because of their higher complexity and higher number of parameters. Research is still new on this, but having a frontier model analyze your code files versus a small, local model for the same task seems to be enormously wasteful. If you must use a frontier model for something, have it do that work after receiving the output from an agent using a small model to read and summarize your code.
Plus how many times do you re-prompt your local model vs Claude Fable or Opus for example to get the desired result?
…Almost never? I’m not a fan of letting AI do much of ANY of my coding, because it will inevitably bloat my codebase with garbage regardless of which model I use. So I severely restrict my model usage to simple, clearly-defined, narrow-scoped tasks that can save me a bit of time, and that’s it. With guardrails and discipline like that, I barely ever have the need to re-prompt.


Are you serious?
I’d love to see some data to back up the assertion that frontier models are somehow cheaper and more efficient than running a model locally.


No, there really isn’t. The frontier models are created through massive plagiarism. They’re designed to be addictive to use. They consume massive amounts of resources to feed you slop. They are inherently unethical. We’re burning the planet down to keep them running, and we don’t even have a demonstrable financial ROI to show for it.
Stop using them. If your employer makes you use them, maliciously comply by wasting tokens until the financial pain is too great for them to bear and they stop. If you yourself are addicted, switch to small, local, open-source, open-weight models you can run yourself. You won’t burn the world down running a small model on your own computer.


That can’t possibly be the actual headline.
[Checks headline]
Well I’ll be damned…


It’s becoming painfully obvious that there is no way to ethically use frontier models powered by these monstrosities. It is currently 100 F in Tuckahoe, the largest city in Henrico County… and they’re asking people to not use electricity so that these heat-and-pollution-generating slop factories can use it instead.
This is insanity.


Doesn’t it work out to something like a full kilometer of the things in order for it to work? The idea is pure madness.


Well… Okay then! This has genuinely been a pleasure.


Hey, before I say anything else, I just wanted to tell you I’ve been enjoying this conversation. It’s nice to be able to disagree with someone on something without it becoming a religious war. :)
I’d say they’re the ones most likely to get screwed on this but I just don’t see how a tech titan drawdown causes the whole world go into a great depression
I don’t think it will cause a great depression, but I do think it will cause a massive recession. If you look at the S&P 500, 35% of it is tied up in AI-related stocks. If AI crashes out, that’s a truly massive hit to the domestic economy, and there are certainly going to be ripple effects throughout the world.
Honestly, the best thing the world could do right now is something a lot of countries are already scrambling to do because of Trump: Decouple their economies from the United States.
Maybe I’m just too optimistic but if captain dickfingers giving the world a tariff hit, a global pandemic and “the worst oil crisis ever” can’t put a dent in things, I just don’t think a pullback on AI spend is going to do it (unless there’s a ton of other structural problems in the global economy that bloggers will point out how obvious it was only once it fails)
Dickfingers! Excellent nickname for the orange sack of shit. But I’ll remind you of the old saying: The market can remain irrational for longer than you can remain solvent. And right now, the market is unbelievably irrational.
The valuations and market caps of these companies are completely disconnected from their profitability (or extreme lack thereof). Among all of them, only NVIDIA seems to be making actual money, and even with them there are some indications of very esoteric math being involved, too. They’re “investing” money into the AI model providers and then having the AI model providers use that money to buy NVIDIA GPUs. Then they book those sales as profit, even though it’s the same money they just invested.
It’s not something that can continue indefinitely. Either the model providers have to show that their unit economics works - by putting actual profit into actual bank accounts - or they will eventually hit a point where no matter what the funny math on their books says, they don’t have the cold, hard cash to pay their bills.
Remember how Uber was about to go broke? Not going to lie, this literally feels like just more circlejerking about companies that are spending a lot and will go broke any day now…
Maybe. But there’s a material difference between Uber / Spotify, and the AI companies. Did you ever look at the SEC filings to go public for either of them? I actually did. They had detailed plans for how to eventually achieve profitability. Uber went public in 2019, and wasn’t profitable until 2023, but they had a solid roadmap for getting to that point, and now they’re consistently profitable.
Our ability to look at profit plans is limited at the moment. The only publicly traded AI company right now is SpaceX, and their filing is… hilarious. Grok will never be profitable. Not remotely. And their filing is full of WeWork-style insanity. They don’t have anything remotely like a roadmap to profitability. Their stock price is 100% speculation. Yet it keeps managing to tick up over $170, at least for the time being.
Speaking of WeWork, I think they’re the model SpaceX, OpenAI, and Anthropic are following. The company raised $12.8 billion in financing, and ultimately reached a valuation $47 billion, mostly from investment by SoftBank - the same bank funding a lot of AI companies now, and which owns 11% of OpenAI.
But it never made profit. It never had a roadmap for profit. It never had any means of bringing in income higher than its operating costs. WeWork declared bankruptcy a few years ago.
Reddit and Lemmy were right about WeWork. So the question now is if the AI model providers are as economically unviable as WeWork always was, or if there’s somehow a path to profitability like Uber or Spotify. SpaceX’s filing doesn’t fill me with hope on that, and the fact that both OpenAI and Anthropic are delaying their own moves to go public doesn’t fill me with hope, either. It’s not the behavior of a company with unit economics that work.
Side note: None of this is to say that Claude Code or Codex or any of their coding tools are bad. I just don’t see how they can operate them profitably. If you’ve got a good product, but the only way to get companies to adopt it is to sell it at a loss, you will eventually fold.
Sorry if I hold a grudge and you are not as dim-witted and not good with computers as the average lemmy user, it’s hard to shake hearing the same prophecies I’ve been hearing about other high spending companies for decades
That’s fine. It’s actually why I mentioned WeWork as a counter-example. Because you’re right, the group-think on both sites (Reddit and Lemmy) can blind people.
On inference alone, both companies project profitability
https://aiafterhours.substack.com/p/openai-vs-anthropic-the-121-billion
https://www.tradingkey.com/analysis/stocks/us-stocks/261756528-anthropic-openai-ipo-tradingkey
That’s true. They project it. Using non-GAAP accounting, and without letting anyone know in detail how much computing for inference costs. Their claims resemble WeWork’s claims, pre-bankruptcy. Bluntly, I don’t believe them.
And I really have to emphasize this: Focusing on the cost of computing for inference all by itself and excluding all the other costs of the business is just crazy, even if inference when looked at by itself can be theoretically profitable.
They are very clearly in the business segment, I’ve heard nothing like this for Qwen or GLM or local hosted models, in fact when self hosting was bought up the dev’s mentioned “??? you can’t self host claude ??”
No argument on Claude being in the business segment, that is absolutely true. But at my company at least (again, this is admittedly an anecdote), the skyrocketing cost of tokens has us working on implementing local models and models on the network edge. We’re also severely restricting token budgets and having devs do as much as they can by hand.
Maybe your company hasn’t reached the point where tokenmaxxing with Claude is frowned on, but the costs are enormous. And the thing is, they have to be enormous for Anthropic to have any hope of ever recouping their losses. It’s not like Claude is a loss-leader, it’s their only product.
This is like saying I work with actual system admins, they say that Windows is terrible and that it’ll be the year of the linux desktop any day now
You need to vibe check the office and devs vs the engineers
The vibe in the office (and what I’m reading) is that claude is gold
I do. I am one of the devs in the office. My boss currently loves me, because my token cost is $0 and I still get my code written. I do most of it by hand, and some of it with Qwen running on my local system (I have a workstation with a good enough GPU to run it). I have it wired up so GitHub Copilot uses it. And I’ve been teaching other developers how to do it. Once our current hardware refresh cycle is complete, our token budget is predicted to drop to almost nothing. We’ll probably still use Claude here and there, but the bulk of our work will be doable without it.
In terms of quality, Qwen 3.6 is at about the same level Claude Opus was a few months ago. I don’t see how Anthropic can compete with that over the long term.
That said I spoke to a normal person the other day and they hadn’t heard of claude at all, blew my mind
That doesn’t actually surprise me. Claude is still a niche product when it comes to general consumers. As far as the public is concerned, AI = ChatGPT (and the annoying Google AI summary).


Yeah, that’s the one that stands out the most to me. I had no idea Origa had died. :'(
The problem is that synthetic data is not fit for that purpose. The more of it you use, the worse at dealing with the cases LLMs get.
Think of it like this… You feed a language model a bunch of genuine human-written content. Great. Now it can produce the most likely text in a lot of cases. Word combinations that rarely appear in written language rarely get generated, so most of its synthetic data lacks those rare - but still valid - combinations.
Train it on this synthetic data, and now more outliers and rare combinations get filed off. Rinse and repeat.