Yea I really resonated with the messaging and mission of Migadu. Super reliable hosting. Have had it for nearly 5 years now. (of course now that I recommend it, I bet they're going to have an outage...)
Metronome has been experimenting with real-time payment on usage based billing for tokens. That technology leverages micro payments using crypto. It's cool but all it's doing is moving the accumulated balance from a provider onto your crypto wallet.
I mean one of those makes them look better than the other. That makes it tricky to be able to tell what is truth vs half truth vs a lie. You're right though, based off of what they claim, their great success is the cause of their failure.
How hard is it to up scale their servers though...
I was going to interject that very sentiment. The upfront costs of running any model of value to the average consumer is so large to be basically non-existent. We need hardware to become commoditized and capable of handling LLMs at the scale necessary to enjoy these agentic workflows.
Mmm, software too though. Even those with capable hardware (Macbook Pro for instance) are usually incapable of getting over the hump of assembling an inference framework, agentic harness etc. If that ever becomes easy, watch out! Governments will not appreciate the user/developer dichotomy being blown away.
Your analogy falls apart somewhere for me but I can't pinpoint the exact correction to make your analogy work for me. Either way I agree. If they don't like only being able to buy digital PS games via Sony's platform, go to one where they can.
That's not the fight I'd personally focus on. I'd personally focus on them revoking digital access after you've "purchased it". Then again, I don't have a game console and I don't "buy" digital media. I either rent via streaming subscription or own the physical.
They make games, they just don’t make all the games. Similar like a grocery store having their own brands in the shelves, you don’t expect to be able to buy these at other shops.
And yes the upfront costs being subsidized is exactly the point.
The PS5 was breaking even less than 6 months after its release. They don't subsidize anything. BTW, Nintendo also don't subsidize it. The BOM of the Switch 1 and 2 is lower than the total price including the margins for retailers!
Shops and distributors take a big chunk. The shoppe price is not given in whole to manufacturer if it was who would be paying the shoppe rental and employees ?
I can see why you might think that, but that is not what this article claims or what Sony said.
The CFO said this in 2021:
“But the profitability of PS5 in -- there was this view of reaching breakeven point in June, but PS5 breakeven point for standard edition, more correctly on standard edition, that's how we explain the situation. And it's been progressing according to the plan. Overall, hardware profitability and peripherals hardware profitability, including peripherals, as we have been saying, is proceeding smoothly.”
And PCMag is inferring this means ONLY the $499 edition of the console which they previously suggested could become profitable later. It doesn't mean every edition of the PS5, and things have changed since 2021.
And Sony said it was having losses on consoles as recently as 2024.
As far as I can tell, despite Sony saying they would like the consoles to be profitable, they are still subsidized as of right now in aggregate. Perhaps some models are more profitable than others.
I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally.
Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.
I compare it to the 1 thousand monkeys on a typewriter. In this case it's 1,000 monkeys with stale training data of everything ever written and the ability to search the web.
I agree with this take in a lot of ways. If you slash the token cost and increase speed for each token 1,000x, who cares if it takes even 20x as many tokens to achieve the goal?
And also, there are lots of tasks where models today are fine with doing. If you think of these things like appliances, who cares if it's not quite as powerful as the next generation? It was purchased to do a task, it still does that task very well. It feels like being in the 90s and asking "why buy a server today when they're going to be faster next year? Just keep renting mainframe time." Well maybe I just need a box to run our HR and payroll system, and this box manages to run it fine today.
I would hire malleable interns fit for replication, and teach them all the things they would need to know. But for some reason I was forbidden to do that, so I'm forced to settle for cheap LLMs instead.
This is my thought as well. Models have to be intentional about which tokens they burn because there's a real lag time. If you can just fork out 10 different reasoning sessions at once with no regard for token waste/lag, you can compensate a smaller model with just doing more at once with it. No idea if this is reasonably true though.
I think that only works if you have checkpoints where all that reasoning can be checked against reality. Otherwise you get an army of armchair experts. LLMs are hilariously bad at home improvement advice btw, where reasoning alone won’t get you far.