“It’s very hard to collect data that trains the model to be better than humans, because there are very few humans who can create that data,” Raj said. “You want to make it better than a Fields Medalist or a Nobel Prize winner. How do you collect that?”

While there, he worked alongside other researchers to understand the failure points of models hosted on cloud computing infrastructure and offer specific solutions, he said. He also created “synthetic” data that, unlike human-created writing or code, is generated artificially before being reused as training material for the LLMs, Raj said.

Creating synthetic data that matches the quality of human-created data is a tall task, Raj said. Computing power can be acquired relatively easily, but finding the data to support the project is uncharted territory, he added.

sigh

Garbage in garbage out, even if the garbage is synthetic that doesn’t make it not garbage…?

  • subignition@fedia.io
    link
    fedilink
    arrow-up
    2
    ·
    10 hours ago

    Wow, it’s almost like that’s an obvious and fundamental limitation of the technology or something.

    It’s darkly funny that they are plainly giving up the game here while seemingly not getting it at all. Of course LLMs can’t be better than the data they’re trained to imitate in the first place. That’s not news. Just about any layperson who’s experienced science fiction in the last fifty years could have told you: If you want to create something more intelligent than humans, you need some mechanism of introspection that allows the system to refine and improve itself with minimal outside intervention. Some kind of artificially created… what was the word again… oh, yeah. Intelligence.

    • supersquirrel@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      9 hours ago

      Well said.

      The synthetic data thing here pisses me off too, synthetic data has uses in science predicting sensor responses, comparing reality to expected findings, and analyizing related phenomena to the artificial data at a fine resolution with computer modelling but none of those things have to do with establishing a ground truth for what the model considers part of reality, part of its understanding of reality or part of the facts that supposedly underpin the reality.

      If a scientist wants to analyze several types of algorithms and compare them maybe they might make a set of synthetic data that is artificially clean and simplified in order to compare and contrast the behavior of the algorithms especially at their edges and extremes. Note however that nothing about this process makes the algorithms smarter, the generative part is what the human scientist learns by observing what happens when the synthetic data is inputted into an algorithm. You need a human brain that understands context, understands the limits of a model vs the rest of reality, and understands things that aren’t explicitly said about the framing context of what is being examined.

      “A.I.” is a lossy data compression algorithm, there is a fundamental “knowledge entropy” here where the end result can never be smarter than the raw data because the “A.I.” can do nothing but apply a lossy data compression algorithm to the training data.

      This is not a cynical take on the potential for artificial intelligence but rather a hopeful and heartfelt thanks to the professions of librarians and archivists, for surely it is the curation of a quality data set where the genesis of intelligence happens. If nothing else Machine Learning proves that with brute force…