<- Back
Comments (464)
- podgorniyJust couple years ago mr altman was promising open AI, to benefit humanity. They even forgot to remove this part about "openess" from the company name.Today chineese deliver that promise and usa people freak out like they have any skin in this game. Enjoy the ride leader of the free world....
- WinstonSmith84I'm rather scared of US models - if Anthropic was the only AI provider in the world, it's easy to see that common people would have no access at all. Thankfully there is OpenAI which compete 1:1 with Anthropic (at a slightly lower cost) but most importantly the Chinese models keep Anthropic, but also OpenAI in checks.And I'm saying this as someone working for American companies.
- throwaw12Lets do "who's afraid of US models" version:* Me, as an individual, because I might not be able to pay price hikes, because my revenue (salary) is much lower than what they want and I can't support my expenses via huge bank loans.* Again, me as a new entrant to the industry, LLMs are basically pay-to-play games, again related to price hikes, new entrants might not be able to afford paying those prices 24/7 - which you need when learning new things.* Any non-US company, US can block the models which can disrupt the whole business.* Even some US companies, for example if you operate in EU and EU somewhat changes their mind and follow the ICC and require you to stop working with Netanyahu (war criminal as per ICC), then following laws in EU, might create trouble to your whole business.
- aaronrobinson“ Let the frontier labs win by being better; don’t let them define safety or security, or pull up the ladder of humanity’s collective knowledge”Love this.
- ballon_monkeyThe 2 things people need to remember:1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
- jke_kangPeople seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."
- _aavaa_> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillationSounds great to me; live by the sword, die by the sword.
- kinj28I am afraid — if Chinese models go mainstream it has a clear way of pushing its narrative way beyond its otherwise borders. More like a Trojan horse it is for the Chinese.Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough questionhttps://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...
- tristanjThe people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models for free. If the frontier labs are forced to cut prices and join the race to the bottom in token prices, these valuations are unjustified, and VCs will face enormous (paper) losses.
- OleksandrCThe article makes a point about agent harnesses being sticky (the supposed moat). I have been building my own agent harness for a while, and I can tell with confidence that the harness almost does not matter, the entirety of the AI magic is the model itself. The harness can be almost barebones (like, for example, mini-swe-agent used for benchmarks), and yet the model still does the task just fine.So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
- faangguyindiaI operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.There are also half a dozen other companies from China continuously hammering our clients’ websites.I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing regioncredit:'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopocShenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.Assuming that China only distills is a huge mistake.It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
- wxw> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
- rafmartomI am more afraid of the US models to be fair. A country with no clear direction in many regards , that is threatening day in and day out the rest of the World for its own interests.
- benrutterI loved this article! Regardless of how you feel about AI as an industry or tool, the economics of AI is fascinating. It's awesome to see something like this that gets into the business side a bit more.I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?
- Talpur1I personally believe opensource or may be state owned LLMs are future, every country on earth should have its own national LLM,trained on country's own data, and then allow its public to use it for free
- jadamczykI think releasing models for free is some 4d chess move by the chinese. Big chunk of the us stock market is fueled by ai mania, if the frontier labs turn out to be drastically less valuable than first believed, the downturm may be very bad. Think of all the big tech companies that have a ton of debt that they took to pour money into AI. It seems like a similar tactic to what Chinese car manufacturers are doing in Europe but the result may be more dramatic.
- bg24I think in general rest of the world needs to take notice (not saying afraid), starting with the US. It cannot be taken for granted that China's frontier labs will be a few months behind. They might be at par or exceed.The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
- daitangioI think the most interesting part of the article is the Huggingface incident at the end:>Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!I understand some guardrails are needed, but it is becoming increasing problematic manage them without a strong public discussion.
- credit_guyPeople who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.3. The frontier labs are also investing more and more in building an ecosystem around their models.I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.[1] https://www.anthropic.com/careers/jobs
- gagandeeprangi1Thinking machines will do it for usa
- throwa356262According to openAI's own @deanwball: Even OpenAI isn't buying this distillation talk:https://xcancel.com/deanwball/status/2078133895766114412#m
- 0x38BExcellent article; the argument towards the end for allowing distillation for US companies is compelling:> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
- vachinaNothing changed for China. The only difference between then and now is that now China is the one selling the products, instead of western capitalists taking a cut off the COGS and selling price.You techbros need to get off your ass and go to work.
- oeziThe article makes a great point that the token industry is going to be commoditized as time goes on.Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.
- deaux> This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
- namar0x0309I haven't had the time to look into recent and past history, but my intuition points at failed empires having similar "elites" starting to stagnate innovation for the sake of "protectionism" - whatever that means.
- abhinaiThe author doesn't seem to realize that a healthy margin has been built into the inference pricing. Once low cost open source inference providers get their hands on powerful frontier level models, there would be a severe margin compression for OpenAI and Anthropic.Source: https://martinalderson.com/posts/the-upcoming-ai-margin-coll...
- tesnorindianAre the open weights models accelerating layoffs in Indian IT sector?
- spenvo"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence"That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
- atmosxNo one should be afraid of anything. Fear is a terrible advisor. Keep your eyes open, try to read the context as careful as you can and adapt as best as you can. Don’t spent too much time trying to be an oracle, never works out…
- oska> Second, intelligence isn’t in fact a perfect commodityWe have got very far from Cicero's coining of the word 'intelligentia' (from inter legere, a 'reading between' and hence discernment) when people talk about 'intelligence' as a commodityPeople have been decrying the 'cheapening' of the word intelligence for over a century now, going back to Psychology's adoption of the word and coining of nonsenses like "Intelligence Quotient". "Artificial Intelligence" is just the latest degradation of the original humanistic meaning, and now people aren't ever bothering to prepend 'artificial' to their idiotic use of the word
- simonreiffI fully agree with everything in this essay. Make distillation fair use. And let us use Mythos/Fable and Sol and successor or future models for all cybersecurity purposes.
- zkmonWhat's wrong if the roles of USA and China are reversed in technology? Why does the rest of the world care? It's not as if USA has done a great good for the world, and China has evil intentions towards the world. Infact it is the opposite in the case of AI so far.
- mattas"Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute."It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
- adrienfr31I am French, and I am sad to see that French models aren't being talked about and that it remains a battle between the Americans and the Chinese.
- softwaredoug> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.But that's a big if we just don't know for sure.1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...
- ArtRichardsNew, smaller models can outperform the previous generation's foundation models.What if there's a way to extract the commodity of intelligence from smaller models?I've seen for many use cases it's well enough. :)
- golly_ned> I expect the inference market to grow much faster than training costsThis was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
- softwaredougHaven’t we been in this “China is 3-6 months behind” for a while now (maybe up to a year? Longer?)The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI storyThen Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
- overfeed> [Anthropic/OpenAI] are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs. Second, intelligence isn’t in fact a perfect commodity, in part because applied intelligence makes itself smarterIs he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
- hexatorI'm worried that any ban on Chinese AI models might be an excuse to get mass surveillance.
- minrawsMe I am, so very afraid of actually decently priced inference.
- ggmA reminder any comment about risk FROM china, invites a "Tu Qoque" facing the other way. The paranoia here is probably fully symmetrical.I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
- nottorpIt's say Anthropic, Allegedly OpenAI...
- coretxThe best model is the model that runs best on your hardware.
- nl> because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?Is this an assertion that is backed by evidence?From the Elon/OpenAI trial:> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
- throwitaway222A company making a decision to allow use of chinese models is a company also choosing to send tons of various credentials to chinese model companies. These will just get scooped up, OpenAI and Anthropic can probably hack into anything at this point if they wanted to.
- ilamontBut it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum.I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
- ab_wahab01Honestly, as someone from a developing country, this shift is good for us. US frontier models are too expensive for us to use regularly. Chinese open-source models/subscriptions are really good to use.
- nunezI really enjoyed reading this.This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
- alizakiThere is no “Chinese LLM”. Each “lab” is distinct and their models behavior is as unique as those from OpenAI and Anthropic
- iLoveOncallAll the article relies on the premise that selling tokens is profitable. I don't see any indication for this, and it makes the whole house of card crumble.
- jonathanstrangeGemini Pro has become so bad for my purposes -- copy & paste Go programming and code analysis with the web GUI -- that I'll take any model with equivalent capabilities at the same price or lower. I don't care where it comes from, I'm not dealing with state secrets and, frankly speaking, US corporations have an abysmal track record regarding safety and surveillance.
- jmclnxOne thing I have not seen mentioned between Chinese AI vs US, population.China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.So I believe, China will end up owing AI.
- fellowniusmonkThe U.S. "executive" class is so obsessed with the "exploit" part of the explore/exploit cycle that it's very clear they are prematurely closing advancement. Better a little money and power for them now than a lot of money and power for their country/humanity.This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
- purplepatrickCommenting wholesale on some folks who are asking for hard evidence. I cannot provide that either but can contribute some empirical data.I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
- NooneAtAll3I don't understand the premise in the beginninghow is running servers supposed to be 0 cost, while running ai inferrence isn't?
- Havocoh wow - hadn't realized they decided to opensource Qwen 3.8 Max. That's pretty big news.
- joshtSomeone (anyone!) get David Sacks on the horn and tell him to read this.
- jdla1oThe future is SLMs and China is going there..
- holodukeSo happy that we have finally 2 countries playing the competitive game. No more secret deals between competitors. No real competition. A race to the bottom is always a good thing for consumers.
- anonundefined
- pupskipperThe fact that Anthropic has a model like Mythos means that counterpart countries like Russia and China are not far behind, if they haven't already developed something similar or better.
- lenerdenatorIt really is amazing that China went from the country that hacked Google out of its market to a trusted source of AI in the tech world.
- zzzeekthe leader of China praised Open Source in a speech. Crazy times
- ChrisArchitectRelated:Ben Thompson is wrong: US frontier labs are right to be panickinghttps://news.ycombinator.com/item?id=48982061
- Alien1BeingThe options are to use LLMs from a country run by a psychopathic regime or alternatively to use Chinese LLMs.
- jdw64While intelligence is said to be a replaceable commodity, oil and copper can be used in nearly the same way even if you change suppliers as long as the quality grade is matched. However, I question whether two models that produce the same benchmark answers are actually interchangeable in real world use.Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
- panchtatvamIs this article written by AI ? It looks so.
- Joel_MckayDistilling models using more advanced LLM is not a new phenomena. It is a cost effective strategy in a highly competitive emerging field.Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
- alfiedotwtfAnswer: US investors
- zuzululuMy thinking is that with the current narratives out of washington we are on track for a ban on Chinese models and possibly sanctions against Chinese AI companiesI think it is the right move to protect American interests
- icasenot enough people
- kdqedI'm honestly more afraid of Claude
- npnwhat a horrible article. full of misinformation and dishonesty.1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is. but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency. and they even plan to build their own inference hardware, too. and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.and you say there is nothing to afraid?
- anonundefined
- marwaneeti think most is vcs
- chewslater secondaries investors in openai/anthropic. It's like time traveling into the spacex ipo.
- kmeisthax> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
- troygentic[flagged]
- andrewdubinsky[dead]
- nttylock[flagged]
- ngl999[dead]
- brandopn[dead]
- sjreeseKellogg School of Business -- he said -- token as a commodity and therefore Open AI is constrained .. ha ha ha hee hee ha .. Well... you build a better mousetrap, and DeepSeek, K3, and ByteDance are just that -- just as good and fit to purpose -- What is needed is to build on top of -- not paniteir (invade privacy and kill people with the information) -- not USMC AI -- use PI's as overwatch killer drones -- but how can I make harder steel, longer-lasting, seawater-resistant concrete, faster time to build housing, better enforcement of USDA rules and FDA adverse enforcement, and better EPA water cleanup, a better FTC for consumer goods -- that is, if I buy an item, that item is safe and built to purpose -- ANYONE not talking about public protection of consumer rights usng AI, is wasting your time
- warshinderPeople with 401k’s and retirees?
- nnmLooks like authored by AI. With quite some reasoning, but no real data to back up main point.
- sharadovWhat makes the Chinese models this good? I don't believe it's distillation alone.This from OpenAi's Head of Strategic Futures "Some observations on Kimi: It's a very good model! I don't think its performance can be explained away by distillation or anything like that"https://x.com/deanwball/status/2078133895766114412China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.To win on the AI front by any means necessary.
- oeatwellI think frontier labs should start building ecosystems by partnering with companies that already have software products, and even collaborating with hardware manufacturers. The ultimate goal should be to create a much broader range of products that integrate naturally into people's everyday lives.The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.