“It’s very hard to collect data that trains the model to be better than humans, because there are very few humans who can create that data,” Raj said. “You want to make it better than a Fields Medalist or a Nobel Prize winner. How do you collect that?”
…
While there, he worked alongside other researchers to understand the failure points of models hosted on cloud computing infrastructure and offer specific solutions, he said. He also created “synthetic” data that, unlike human-created writing or code, is generated artificially before being reused as training material for the LLMs, Raj said.
Creating synthetic data that matches the quality of human-created data is a tall task, Raj said. Computing power can be acquired relatively easily, but finding the data to support the project is uncharted territory, he added.
sigh
Garbage in garbage out, even if the garbage is synthetic that doesn’t make it not garbage…?
Wow, it’s almost like that’s an obvious and fundamental limitation of the technology or something.
It’s darkly funny that they are plainly giving up the game here while seemingly not getting it at all. Of course LLMs can’t be better than the data they’re trained to imitate in the first place. That’s not news. Just about any layperson who’s experienced science fiction in the last fifty years could have told you: If you want to create something more intelligent than humans, you need some mechanism of introspection that allows the system to refine and improve itself with minimal outside intervention. Some kind of artificially created… what was the word again… oh, yeah. Intelligence.
finding the [human created] data to support the project is uncharted territory, he added.
Fortunately for the executives at hyperscaler corporations, they know this isn’t their problem.
They plan to cash out well before the model collapse is revealed, and leave less-wealthy suckers holding the bag.



