I guess this is being down-voted just because it feels pro-AI. Believe me I get it, you can look at my anti-AI post history. But I think it's important to also keep an eye out for real signals towards general intelligence, which is why I wanted to share this. I'm not celebrating it, just bringing awareness.
For as long as it is an LLM I will not call it intelligence for it does not think. It's an approximation or an estimate at best. While capable this is a major hurdle these companies will face to replace skilled workers. We are a generation at least and a ways off from general intelligence, and not algorithmic output for a given input (idempotency).
I do however agree that we would remain informed about the ever changing world around us; this however, is until reviewed by an unbiased third party in masse a marketing document. I will be taking it with a grain of salt, but will take it as an improvement over the prior versions.
Now with that out of the way... interesting. I will be watching this progress, but am far less optimistic in this being as capable as they are suggesting. It's more than likely another Mythos hype attempt that is not revolutionary, but a nice addition to existing systems or augments to staff. Now if we assume its 100% as capable as this suggests then we are in for a ride.
Its all semantics. Im curious, can you define what it means to think? Maybe the term is broadening now that we found some kind of similar process that looks like what we imagine thought to be. AI is bad and evil yadda yadda yadda im not defending it i just really like this philosophically or whatever and itsd be cool to hear why you differentiate "thinking" from human thinking
I agree with most of this. I've also said in the past that LLMs cannot think, and I think that's still true for most models. The reason ARC-AGI-3 is interesting is that it was specifically designed to test reasoning, adaptability, novel problem solving, planning, memory, etc. So it was a surprise to me that Astra was able to defeat it so effectively, and that Astra invents algebras for each novel task.
But I agree we can't trust OpenAI if these results are self-reported, and we may not be able to trust the ARC Prize Foundation fully either. Extraordinary claims require extraordinary evidence, so we need replication, transparency, and proper open science to confirm things.
I also agree with ARC Prize's conclusion, that there are still capabilities any AI system would need to demonstrate before we can claim a full general intelligence.
OP If you don't wanna be downvoted so much you gotta use a more negative framing like this OpenAI announces new slop generator GPT-6 Astra: Unfortunately not any better than the last one
I'm okay with the downvotes. There's a stronger anti-AI sentiment now than even just a month ago, and on the whole I think that's going to be a good thing.
Same, im very interested in the technology, ive followed Machine Learning stuff since before most of the kids using ai to cheat in school were born. But the way things are going, the corpos in control of the cutting edge are making such abhorrent and misanthropic decisions that anything but pure hatred for the use of their tools is unacceptable. I struggle to justify using even local models, i just feel icky now.
Wow they benchmaxxed this pretty quickly. Wasn't this benchmark released only a few months ago with like 0% success rate from all major AI models?
Yes, it was released in March. One important point is that these results are based on the semi-private test set. Perhaps it's best to wait until it is measured against the fully private test set, but that might not happen for months.
Technology
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.