The world's most-powerful AI model suddenly got 'lazier' and 'dumber.' A radical redesign of OpenAI's GPT-4 could be behind the decline in performance.

L4sBot@lemmy.world · 1 year ago

The world's most-powerful AI model suddenly got 'lazier' and 'dumber.' A radical redesign of OpenAI's GPT-4 could be behind the decline in performance.

nbailey@lemmy.ca · edit-2 1 year ago

The model has become inbred because it’s now impossible to scrape the web without AI content getting ingested, which is full of “hallucinations” and other weird artifacts. The last opportunity to get “uncontaminated” training data was sometime in mid 2022.

Not to say that it’s causing this particular problem, but this issue will emerge eventually. Garbage in = garbage out. Eventually GPT-19 will grow a mighty Habsburg chin.

TheGod@lemmy.world · 1 year ago

Removed by mod

Chaotic Entropy@feddit.uk · 1 year ago

All the articles with very specific titles, but then incredibly generic content, piss me off to no end.

NotYourSocialWorker@feddit.nu · 1 year ago

Part of the reason why debugging windows is such a pain. Another part is the so called experts in the forums.

Chaotic Entropy@feddit.uk · 1 year ago

Make sure your drivers are up to date! Another job well done.

Ultraviolet@lemmy.world · 1 year ago

Also the articles that are plagiarized but run through a thesaurus bot to bypass search engine penalties for being plagiarized, often to the point of incomprehensibility. Yes, I’d love to read an article about my favorite vagabondlike, Deceased Cells.

jantin@lemmy.world · edit-2 1 year ago

Maybe not yet, but…

Spez will turn Reddit into a bot farm and sell this as training data
Musk turns Twitter into a bigoted cesspool and will sell this as training data, which will subsequently be flagged for low quality (also: a botfarm)
Threads is a corporate ad dashboard (and we already know how easy it is to GPT copy) and Zuck will sell this as training data
Facebook is either dead or only good for boomers and Poles
blogs are dead
Fediverse is out there waiting to be scraped but possibly too small to sustain a big model

We’te getting there, hopefully.

cyberpunk007@lemmy.world · 1 year ago

Scrapped?.. Or scraped?

jantin@lemmy.world · 1 year ago

absolutely scraped, fixed

damnYouSun@sh.itjust.works · 1 year ago

Also We'te, which I believe is a Klingon name.

cybersandwich@lemmy.world · 1 year ago

…is Facebook popular with Polish people? Or was this a weird polish joke I don’t get?

AccidentalLemming@lemmy.world · 1 year ago

You’re saying that as if bot-generated content farms haven’t been around since forever.

But you’re right that now it’s easier than ever to publish low-quality articles that are indistinguishable from well-researched ones.

RIotingPacifist@lemmy.world · 1 year ago

Nah GPT makes it a lot easier, it’s the thing it’s actually good at.

Before they were autogenerated with bad English, GPT can generate good English that is equally devoid of content

TitanLaGrange@lemmy.world · 1 year ago

I suspect future models are going to have to put some more focus on learning using techniques more like what humans use, and on cognition.

Like, compared to a human these language models need very large quantities of text input. When humans are first learning language they get lots of visual input along with language input, and can test their understanding with trial-and-error feedback from other intelligent actors. I wonder if perhaps those factors greatly increase the rate at which understanding develops.

Also, humans tend to cogitate on inputs while ingesting them during learning. So if the information in new inputs disagrees with current understanding, those inputs are less likely to affect current understanding (there’s a whole ‘how to change your mind’ thing here that is necessary for people to use, but if we’re training a model on curated data that’s probably less important for early model training).

I don’t know details of how model training works, but it would be interesting to know if anyone is using a progressive learning technique where the model that is being trained is used to judge new training data before it is used as a training input to update the model’s weights. That would be kind of like how children learn by starting with very simple words and syntax and building up conceptual understanding gradually. I’d assume so, since it’s an obvious idea, but I haven’t heard about it.

minorninth@lemmy.world · 1 year ago

That hasn’t happened yet. Most likely they quantized GPT-4 more. It’s still based on the same training data.