Hacker Newsnew | past | comments | ask | show | jobs | submit | adam_arthur's commentslogin

There are an enormous number of tasks that can get by on good enough.

If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.

And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.

Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.

5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?

The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.

I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.


That can all be true but the frontier models will still have a huge market. You're thinking of all tasks as a fixed pie. The top 1% of intelligence opens up a whole new pie, stuff nobody does today because it's too expensive: daily cancer scans instead of one every few years, asteroid mining missions that need ten thousand PhD-hours of planning, custom drugs designed for your specific tumor, a personal lawyer and doctor for every person on earth, auditing every line of code in every bank and hospital continuously and so on.

> daily cancer scans instead of one every few years

Silly nitpick: the reason we don’t do daily cancer scans isn’t the cost, it’s the false positive rate. Invasive procedures like biopsies come with complications like infections that happen at a higher rate and do more damage than the cancer that doesn’t even exist. This dilemma is pervasive in medicine, because our tests aren’t perfect but the thing they’re testing for is rare.


The whole paradigm changes, though, when you can do daily cancer scans. You don't get a biopsy when the scan shows a lump. You get a biopsy after a couple weeks of daily scans showing the lump growing. Plus, having all the data from the daily scans improves your testing accuracy so false positives and negatives are more rare.

The errors are correlated, not random. Lumps are usually benign cysts. If you start cutting people open for every cyst on a scan, you’d kill a lot more people than you’d save.

This isn’t something you can solve with more scans because the tests test for data that is indistinguishable. They look the same on a scan, there’s an overlap in the assay with some random protein with the same binding sites that is only present in 1% of the population, the coding gene in one person gets repeated in a noncoding region in another, and so on. The “more data” that works is a doctor applying professional judgement (which they’re also famously bad at because biology is a fickle mistress).


If you're doing non-redundant tests and your uncertainty bars aren't shrinking, it's usually a skill issue.

If it looks like a duck, it might be a duck - or a painting of one. If it looks like a duck, swims like a duck, and quacks like a duck? The joint duck estimation is much more confident now. There might be a few more observational tests one should administer before committing to a duckhood decision, but each tests pins down variables and rejects confounders. Uncertainties are cut down, and we get closer to crossing the threshold between "duck-informative" and "duck-actionable".

Thus, it's often worth it to improve observability. If you managed to make a certain test more reliable, or cheaper to administer, or reduced the chance of adverse effects? Or, in other words, improved SNR, reduced costs, and reduced costs? You can get more information for your buck. Paired with good knowledge: you can make better decisions more easily.

The fact that the thought of "having more information might be bad actually" even occurs in the field of medicine shows just how far it is from being optimal. Having more information isn't always beneficial - some information is genuinely redundant. Some information is not worth the effort of gathering and integrating it. But if you get more information and it results in worse outcomes? You're doing something wrong.


More data (daily scans) can mean we get better at medicine, though, and more accurate. You're assuming "daily cancer scans" look just like they do today, rather than eg unobtrusive devices in our daily lives that measure changes over time.

The frequent "muh false positives" comment we hear from doctors appears to be a lack of imagination?


It comes from experience.

Yes, agree that token consumption will increase exponentially for the next while.

Disagree that the frontier model is where the economic gains will be realized.

The smaller the relative gap between frontier and non-frontier/open weights, the less pricing power.

This gap has shown only to shrink over time, not expand.

Businesses will pay more for frontier, but not meaningfully more to justify the economics. It's always going to be a low margin business, perhaps outside of cyber security, warfare/intelligence and perhaps drug discovery.

Though the expensive and time consuming part of drugs is doing the trials and getting approval, not coming up with ideas


Sounds kind of like another K shaped economy. Low end models will be highly competitive and low profit. Problems that can be solved by low end models will be highly competitive and low profit too.

Where the interesting work will be is at the median point where cheap models do almost all of it but need to hand off some parts to the SOTA/more expensive models. Seems like there's money to be made by maximizing low end use while maintaining quality.


Certainly there's still a business there, I'm not saying they won't exist. But it's not going to be a monopoly-esque business with so many players in the ring, OpenAI, Anthropic, Google, Meta, Deepseek, Alibaba, GLM, Kimi etc. It will be cutthroat and a race to the bottom on price. And the difference from today -> 6 months ago intelligence will not be very meaningful.

Investors are largely treating these as future monopolies though.

We can already do so much with existing models. Harness improvements are probably more meaningful at this point.

e.g. say most image recognition can get saturated by a model of size xB parameters, so your tool for that can handoff to a smaller model. Document text extraction can use a model of size yB parameters. A model of size zB for summarizing text.

We are starting to get to a point where you can reasonably scope out an upper bound of required size/effort for many common tasks, and if you string these together, the frontier will largely act as an intelligent invoker of more efficient models.

Up until now there have been meaningful gains to each of those types of workstreams by using newer models, but that is starting to no longer be the case.

Yes, I do believe token consumption will rise exponentially from here in the near term. But cost of switching is low, and substantial profitability will be difficult.


how much would you pay for a prompt that could cure cancer? if you're a pharma company you would pay millions to get there days faster than your competitor. as intelligence rises the marginal value it can deliver rises with it.

Something like curing cancer (more realistically, curing a specific kind of cancer) has to interact with much slower real-world processes. The most expensive part of drug development is Phase 3 clinical trials in humans. Even the smartest model in the world can't accelerate that meaningfully. Even much earlier when drugs are just testing in cell cultures, it's a lot slower to run lab tests than to run software tests or mathematical proof checkers.

Or to put it another way, there's enough natural variation in real-world bottlenecks that no pharma company can assume they'll beat competitors to market by using a smarter model.

A really smart model could significantly improve the pharma business if it could identify promising approaches to cancer treatment that are less likely to fail in clinical trials, but I don't think that the frontier labs have data to make that work yet. Much of the biomedical literature is poorly reproducible ("replication crisis") and much of the drug-development-specific data is proprietary, never published in the first place.

I do have hopes that general laboratory automation will go faster with LLM assistance, even if all the LLM does is write Python glue scripts to enable custom workflows and instrument integrations.


Wouldn't daily cancer scans give you cancer from all the radiation? I thought that's why the doc hides behind a wall when giving an X-ray.

yea you see no body are vibecoding games before opus 5 and astra, after they are released games basically got commoditized

If the output is commoditized, how much can you afford to pay for the input?

Is this sarcasm? I assume most studios are integrating AI into their workflows, but I still haven't seen a single vibecoded game that looks interesting.

I was yelled at repeatedly here that AI is not used in gaming.

It seems obvious that labor displacement due to AI will create labor surpluses in existing fields.

Which will feel quite bad for many of us in the present, but net-net be good for society in the long run.

Wages will go down, but cost of goods and services will decline even more.

Markets will be more competitive than ever, which limits the moat/margin capability of many businesses.

It's only really particularly bad for those earning very high wages where there likely won't be any way to materially fill the gap.


What is obvious about that? What jobs are being created long term that can be filled by those displaced? To me it seems like even jobs that are not knowledge jobs are going to utilize ai to cut hiring, not grow it

When wages decline, cost of producing the service/good does down too. And when the costs go down, relative demand goes up.

When relative demand goes up, the total labor needed expands.

How many Uber rides would you take at $5 vs $20?

How many times would you eat out at $10 vs $50?

Would you pay for a house cleaner at $25/visit over $150/visit?

A live in nurse at 50k/year vs 200k?

The costs of all these things will decline materially when the labor pool expands meaningfully due to displacement.

Technology has always been deflationary in the long run throughout history.


Why do you think it will be net-good for society in the long run?

If wages go down, worker power goes down; and if dependence on labour goes down, you literally move closer to the situation of the extraction economies in things like petro-states, leading to oligarchy instead of democracy.

It could be the end of ordinary people's power over society rather than anything even slightly good.


Nominal wage doesn't really matter, just real wages+purchasing power.

Goods and services prices will fall far more than wages for most people.

Why?

Because there will be far more efficiency and competition than pre-AI. It's plainly evident from the nature of the technology.

If one business can do something cheaper, other businesses will do it cheaper too, and have to cut prices to compete.

But the shock will be sudden so it won't feel good for our gen. Future generations will benefit without the drama.

Don't get me wrong, there will be big losers in some fields, and the short term will be painful due to retraining and loss of purpose/emotional toll.

Many of us built careers doing things that may not be relevant anymore. Or relevant in a different, perhaps diminished way.

But the same has happened to many professions throughout history and it's always led to general improvement of the broader public's welfare.


I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.

I think some combination of:

1) Using 1 thread for everything

2) Reviving old threads which are no longer in cache

3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.

4) Non-coding workflow which is more output than input heavy

5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR

Just a guess. I think 3 is likely the primary reason.

You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME


That's my experience too. I've found OpenAI really quite generous with tokens. I sometimes wonder how some people manage to run out of them really. Do they just type prompts that much faster than me or use the highest reasoning mode for everything just because they can? Idk.

I generally agree with those reasons, although using a single thread may be less of an issue than it seems because of context compacting which should happen automatically when you're near the limit.


My use cases are iterative and sometimes require reading a lot of code or reevaluating work.

Token efficiency is near meaningless when the workload is input-heavy. It can't always just choose to read less, depending on the task.

I can have cheaper agents do the reading but it's not appropriate for all use cases because they'll misjudge and choose the wrong things to emphasize, summarize, extract for the bigger model.


Many threads.

I use new threads if relevant old one is uncached. (Often using a skill or doc for handoff instead of requiring full context gathering again.)

I get involved in architecture and specific implementation direction. The codebase is 8 years old and mostly handwritten.

Mostly coding. Some QA.

No cron/CI agents.


3.

The problem is also that at a lot of code bases are not designed around LLMs and their token usage.

If you have a monolete codebase, you need to clearly define in the prompt what modules are involved. And even then, your wasting tokens with the first 1 or 2 steps where it needs to located the modules.

if you have a github codebase, with a lot of your code into nice little repos, its even worse because then the model needs to pull data, and a ton of more steps.

Most people do not open their coding agents in the module directories because "it may need something out of it".

I mean, we used to program by creating utils directories to deal with repetitive code but models (a) find it and use it (but it cost steps and tokens to read), (b) do not realize it there and make their own version of whatever or (c) combination of both.

And ironically, i feel like we need to give up on this idea of reusable code, and literally keep things into single modules, with as minimum external dependencies. That in return reduces searches and thus steps/tokens burn. But very few agents / harnesses have proper implementation of groups/projects and sub-module structures. Aka they only open a single dir, so your then forced to create dozens of tabs > per dir > cli ...

2. That is also a issue if you open multiple agents. Maybe now your working in A, B but C, D, E are not doing anything. And their cache expires... Now you go back to D because A, B needed to be done, and now your paying Cache Write + Input cost.

Its hard to have a good flow to keep things cached, when to really /new and when to not have it expire (and that assumes there are no issue with the provider moving your session around and forcing new cache hits. MiMo did that a lot in the past).

Something that i also advice more and more to people. Get a microphone, download openwisper and talk (text to prompt). You tend to give more information vocally, then writing as its in our habit of programming to not be verbose. Its like people are afraid of long prompts. While just talking to the LLM tend to give it much more information to work with, often resulting it being able to skip steps.


Codex (5.6 Sol) is extremely to the point and direct in its responses.

Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.

No idea why people are still using Claude models.

My impression is they started on Claude Code and never tried Codex or other harness+model combos


I was on the Anthropic train for 2 years, and I tended not to jump between vendors much because it felt like a lot of distraction for little gain. But last week I switched to 5.6 Sol during an Anthropic outage and it made me realize how frustrated I was with Opus 5 and I haven't switched back.


That is true. On 5.6 I have noticed it to be a little bit more like Claude, which I dislike. But still way better than Claude.


I see the complete opposite becoming the case.

In a world where code can be generated rapidly, it's super critical that you have a core few set of people who really understand the macro design of the codebase and can continue to factor it well and iterate quickly.

Adding more people and contributors just increases the probability that nobody really understands the structure of the codebase, it degrades into DRY and unfactored slop.

The cost of reviewing other people's code is almost too high to be worthwhile now... It's much easier to just cut them out and do it yourself.

A core set of very skilled people can just implement whatever change you are doing, but better, cleaner and faster.

I built a fairly large and complex project with Codex and had to spend about 50% of the time factoring things down as I went into well contained modules, had a full understanding of the architecture at a high level. It would have been pretty difficult to do this if bringing in other contributors.

Too many people comment about AI from the perspective of throwing feature A or B over the wall at the workplace, but anybody who has built a huge project from scratch will see how important good design is in regard to iteration speed and result quality.

That being said, there are still areas where changes should be sized reasonably and human reviewed e.g. foundational or very mature software


That makes sense from an engineering perspective, but from a business perspective you don't want the understanding to live in the heads of a small team of people. Companies own the codebase, they don't own their employees. If losing a single employee means losing the understanding for a significant chunk of the codebase, that's a serious risk. With the senior/junior model where you have a senior engineer architecting the system and a small team implementing that architecture, the understanding lives in the senior engineer's head, but it also lives in the head of the person who implemented it, and their teammates who implemented the parts that interact with it and were present for the discussions probably have enough knowledge to figure it out pretty quickly, so as long as the company doesn't lose the whole team all at once they're okay. If you make the team more productive with AI tools you can have the team do more, which still cuts down the total headcount, if to a lesser degree, but preserves the redundant understanding.


30y is keyed to inflation expectations.

If fed hiked to 5% tomorrow, 30y would invert and yield would go down.

It's not as simple as hikes lead to higher 30y yields.


It's not all inflation expectations, either. The dollar has been strong lately due to elevated oil prices---countries that are short need extra dollars to buy oil, so they often liquidate treasuries to get them.


Additionally recent studies have shown high duration/volumes of cardio increases plaque buildup in the arteries.

"Indeed, many studies have shown a relationship between endurance sports and higher volumes of coronary calcified plaque as determined by computed tomography."

https://pmc.ncbi.nlm.nih.gov/articles/PMC11395881/


Wow, their "high" threshold is 150 minutes / week.

That's really not a lot, it's like 3 days / week of running a slow 10k.


Yes, given all the recent data coming out, it seems to me HIIT is a much better bet for general health purposes than LISS style cardio.

Curious what the consensus will be a few years out.


Define HITT.

I guarantee you that HITT with stupid movements is also not good for you. And you can over train HITT. Source: all the people I used to do CrossFit with that got injured or washed out from over doing it.

As always the answer is to listen to your body and don’t over do it. Running marathons is fine. HITT is fine. Lifting weights is fine. In moderation at an intensity suitable to the individual.


Yes, if you set reasoning to none you can force the granularity of the thinking.

It will actually adhere to your request for e.g. 3 sentences max.

Thinking mode will override any instructions in the prompt (at least for other models in my experience).

Of course this will probably hurt performance, but works great for easy tasks that you know are trivial. Tons of pipeline, image recognition etc use cases where this works well.

I'd be curious to see Qwen 3.8 27B low thinking benchmarks though.


There are tons of core principles that can be learned that largely apply across fields.

One prime example, single source of truth for data/concepts. To be violated only when performance is meaningfully improved (denormalized databases). But when you do so, you should definitely recognize you're opening up out of sync issues for that performance gain.

Though for the majority of code, there is no performance benefit to adding multiple sources of truth. Yet it's the most common error I see re: quality.

The sad thing is that software engineering fundamentals and best practices never became widespread or widely taught in school prior to LLMs


I believe evidence points towards absolute level of fat being more important than % for overall health.

E.g. 20% bodyfat at 250lbs is still a lot of fat.

Of course it's difficult to ever get very high on absolute fat if at 15% or below.


Not sure I follow the logic here, but happy to be proven wrong if I'm misunderstanding. It's not intuitive to me that a 300 lb. person at 20% body fat is generally at the same risk as a 200 lb. person at 30% body fat.


Fat is inherently unhealthy for you beyond some low baseline level.

Additional muscle is positive for health, but only up to some reasonable threshold. There is no health benefit to having very high levels of muscle, and in fact it may be negative for your health at extreme levels. E.g. many bodybuilders have trouble breathing, sleep apnea etc.

Both people in the example have 60lbs of fat. 1lb of muscle doesn't cancel out the negative health effect of 1lb of fat


The claim that fat is inherently unhealthy is probably too broad.

From what I have heard from doctors, there is a clear sequence:

intra-organ fat (really bad) > visceral fat (quite bad) > subcutaneous fat (relatively harmless).

The nasty detail is that some people gain relatively more visceral fat than others. These three fats are correlated, but the correlation isn't the same in every human. Some have better metabolism and don't store as much fat inside as others do.


Looking for some references on these claims, thanks! What you're saying registers as "makes some sense, but where's the evidence?" to me. I'm not seeing the connection where total lbs of fat is inherently worse than higher percentage.


You won’t find any reference for those claims. Did you know that the majority of people who have heart attacks have normal lipid levels?

It’s not the fat that kills. It’s oxidated lipids that kill.


I haven’t heard that as a cause of heart attacks. Is “oxidated lipids” the phrase to look for in research?


Yes. Rather, Oxidated LDL for heart disease specifically.

https://www.sciencedirect.com/science/article/abs/pii/S24518...


Do short people have much lower risk of heart disease?


The mean height for men at 20-29 is 69.2" and 80+ is 67.1" (measured in 2015-2018, the last data I've found) [1], which could be interpreted as shorter people having lower risk of all mortality. One could say that people are just getting taller over time and 80+ y.o. were shorter at their 20-29, but we also have the median height for 20-29 fro 1970s, when 2018's 80+ were 20-29 - it's actually 69.7" [2], men in the US are getting shorter over time, so unless people naturally shrink 2" with age the data points to shorter people having lower risk of death.

1. https://www.cdc.gov/nchs/data/series/sr_03/sr03-046-508.pdf

2. https://www.cdc.gov/nchs/data/ad/ad347.pdf


People do tend to shrink in height quite a bit at advanced age.

https://medlineplus.gov/ency/article/003998.htm

"People typically lose almost one-half inch (about 1 centimeter) every 10 years after age 40. Height loss is even more rapid after age 70. You may lose a total of 1 to 3 inches (2.5 to 7.5 centimeters) in height as you age."

So people DO in fact tend to lose about 2 inches at age 80 versus their younger selves.


The actual studies [1],[2] that tried to measure this came to very different rates of height loss though: 3.6 cm and 6 cm from 40 to 80 in men. This does not look like a serious fact.

1. https://pubmed.ncbi.nlm.nih.gov/10547143/

2. https://www.mdpi.com/2072-6643/15/21/4694 (referencing the FHS there, the numbers are hard to find on the study's website).


Even the low end of 3.6cm is 1.4 inches. Feels like you're being rather...goal post movey.


There is no "low end" , such a huge difference shows that both studies are not reliable.


The difference was 3.6cm in the study that measured 10 more years 30-80. And 5cm in the study that measured 10 fewer years 40-80.

The studies were conducted over different time periods when average nutrition changed, and demographics changed. And the study demographics just vary by location.

Height loss with age is incredibly well documented. Here’s a few more studies.

This is just basic medical textbook information. No one in the world is arguing that this phenomenon doesn’t exist.

https://pubmed.ncbi.nlm.nih.gov/32611342/

https://pubmed.ncbi.nlm.nih.gov/24962157/

https://pubmed.ncbi.nlm.nih.gov/10602346/

https://www.sciencedirect.com/science/article/pii/S266703212...


Now, who is moving goalposts?


No one:

Again, from your OP: "so unless people naturally shrink 2" with age"

Are you denying basic science here, or what? Like, what is your deal?


Read what I wrote, read what you wrote. If you still don't get it - ask a chatbot, though you obviously understand and are just trolling.


I’m also confused.


Short people have fewer incidences of cancer, so I don’t think you can say that lower overall risk of death implies lower risk of heart disease.

As for 2, men in the US are getting shorter because of immigration which adds many confounders.


I am not saying that. I am saying that the theory about absolute amount of fat being bad for health, as stated by the comment you responded, seems to be supported by data.


That’s not the question I ask though. I asked about the specific claim that absolute fat was bad for your heart.

But also short people living longer has so many more possible explanations than less overall fat.


> I asked about the specific claim that absolute fat was bad for your heart.

There is no word "heart" or any to this effect in the message you replied to so I did not understand that was what you were doing. I took your question as an attempt to attack the theory on the basis of prevalence of some heart conditions in shorter people.


No that’s not what I was getting at. The post was edited and originally mentioned absolute amount of fat being bad for your heart.


Some are more common in short people and some are more common in tall people. Hypertension is one of the ones more common in short people and is significantly more common than the rest combined, so your overall chance of having some form of heart disease goes down as you get taller. However, the forms of cardiovascular disease more common in tall people (e.g. atrial fibrillation) are more likely to actually kill you.


Do you have a citation for that? I've had thoughts in a similar direction (specifically around the upper/lower intervention points) but have had no luck finding data.


Its not just absolute amount of fat though you also have to account for intra abdominal fat being much worse metabolically than subcutaneous fat.


Does that have to do with sizes of the heart and few other organs apparently being largely constant regardless of height?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: