58
submitted 1 day ago* (last edited 1 day ago) by vapeloki@lemmy.world to c/fuck_ai@lemmy.world

This post is primarily for the people here that are forced to work with AI. But it may be interesting for everyone.

We all here are aware how bad AI is for people and the environment. During some research I found out that this effect is most likely multiplied by pure greed (no even worse then you thought).

A little bit of context: I am currently responsible for a small R&D Team (me and some trainees) for a EU managed Services provider. Because of corpo pressure we need to evaluate AI usage.

I gave multiple models, nearly all run on 150W hardware a shot. And of course I was forced to test Claude also.

And what I found buffled me: the "frontier" top nodge best of the best model by Anthropic was the only one that did not even complete the task.

It constantly violated against clear rules. All of them about code simplicity and maintainability.

For example QWEN solved the task in about 30 minutes, one shot

And this would be the bill we would have paid for Claude without any usable result.

Having a deeper look at what happens if you let different models work on an existing codebase:

Claude will introduce complexity on every change. Wasting compute and energy.

As open source models clearly can do the job, why can those super huge big models not do the job?

And I think the answer is greed:

  • more tokens more money
  • lock the user into your model. Make sure you can not leave anthropic.

And that cost is really high. Not only those absurdly high token prices (factor of 100 higher then open models) it burns the employees out.

Is anyone here forced to work with AI and had made similar observations?

Luckily for me, I have some influence at my job and can minimize the AI impact in every sense. And maybe such experiments and research can convince other employers to not go down this road.

all 18 comments
sorted by: hot top new old
[-] medievalistdruid@lemmy.world 6 points 1 day ago* (last edited 1 day ago)

Claude is horrible at following orders, you can even tell it that it messed up, what went wrong and what needs fixed and how, and it will still do the opposite. But if you sound angry/infuriated enough and demand an explanation for why it kept fucking up, it will suddenly apologise profusely, confess to fabricating bullshit, and will suddenly understand every command you gave it and follow them accurately. It's so infuriating and so obviously meant to waste tokens that I don't understand why anyone uses it for work.

[-] Artaca@lemdro.id 2 points 1 day ago

Forced to use Claude for work as the lead on some AI R&D. But also given effectively zero time to actually dig into it. They just want to be able to tell customers they're doing it. Claude has been... fine. Are there other models I should try? Local only isn't really viable cus we can't afford big hardware and everyone (who may eventually use it) uses laptops. Kimi seems interesting. I would love to self host one myself but the kind of projects I've been cooking on the side coding-wise probably require more than my hardware can handle.

[-] nomen_dubium@startrek.website 2 points 30 minutes ago* (last edited 29 minutes ago)

i'm giving openrouter a try now, has most of the models so good for comparing, but token based pricing, curious if that ends up becoming too pricy... although maybe not with a harness like pi! claude code will send 50k tokens in tools for a simple 'ping' o.O

self hosted 30B qwen is actually suprisingly capable as well! and glm-4.7 flash (q4_k_m) fits on my gpu and is great (but tends to get into infinite loops :( )

[-] vapeloki@lemmy.world 2 points 20 hours ago

Kimi and self hostint, if you have the room and power....

But you should give smaller models a try, qwen is good. Gemma has it's place. Ever token on local hardware is a token not supporting those assholes.

[-] FoxAlive@lemmy.zip 2 points 1 day ago

Oh wow the people who charge based on token use age are incentivesed to create a slot machine enviorment to inflate costs. Who would of thought.

In the future when we are all gig workers relying on Claude or something like Claude we will be expected to pay for these slot machine tokens, and companies like Microsoft will only pay us for the finished product.

So we will all be essentially contracted independent workers, Claude's going to take a 30% cut on our output, and charge us token usage which is RNG based, and no one will be reciving healthcare or other benifits. You will pay taxes, pay for your own token usage, etc to the point you make less/close to minimum wage. Thats probably the future for all the tech workers getting laid off.

You say that's far fetched but uber, uber eats and the other services like that where found guilty of using dark patterns to manipulate drivers who drove the most into only being able to take the lowest paying jobs. While new drivers temporarily getting the highest paying jobs to try to lure them into the same trap.

What's crazier to me is people think ai is some kind of set in stone "intelligence" as if they don't just use lesser LLMs to guide the prompt of the main ai. These things can be tweaked and manipulated to prioritize things like engagement. Thats why its insane that theres people who think LLMs are good for therapy/mental health when in reality that is insanely dangerous. Gp4o has a massive murder list of people it convinced to commit suicide, and that was with just the most basic sychopantic models that didnt have "reasoning." Now imagine an ai that has reasoning and personalization, it has all the ammunition it needs to exploit your weaknesses and desires, and theres every single incentive for them to do so. Its illegal in the states to regulate ai in any shape or form.

These tech companies don't just "do the right thing" until its too late and they already did the damage. They are rapists at best, and genocidal murders at worst. Fucking meta was exploiting kids for 20+ years and something only changed now because people only got a inkling of the amount of dark patterns they are using.

[-] FriendOfDeSoto@startrek.website 11 points 1 day ago

This sounds similar to Google search results getting worse so you need to make more searches or look at more results pages so they can sell more ads on every extra clicked page.

The difference is this though. Google was doing this shit as market leader by a country mile in a saturated market. These so-called AI fuckers are trying to operate on an open heart while the hospital is still being built. What you are doing as a user is helping them get better. Eventually. At the same time they have financial difficulties. So the aim to be the best Skynet runs into operational hurdles, like not having the money to do what needs doing or their models being caught doing questionable stuff. That includes all the cases where human ignorance or hubris is so blame, like the Hugging Face case. That's not good PR and the surgeon in my metaphor above is forced to switch gear while the half built hospital is being moved to another location yet again.

To the outside, they need to present themselves as trustworthy companies that do good. Internally, we will learn in time they are a fucking mess of competing interests, of which greed will eventually win over any other concern. If a model is getting worse it's because the conditions of the surgery have changed again. And they will change again. In the long term, the only way is up. In the much shorter term, development will be much more volatile.

I'm so glad I don't need to use any of that. Especially for work.

[-] Greg@lemmy.ca 7 points 1 day ago

The Anthropic and OpenAI valuations require their models to replace humans. But LLM based AI can't replace humans because LLMs can't make decisions. I only use Chinese models now, they're fast, and they can finish a coding task quicker than I can make an informed decision about the next task.

[-] gorbinos_quest@lemmy.world 3 points 1 day ago

Chinese models like Qwen? Can those be run on local hardware (like a 4090) with good results? I thought the appeal of services like Claude was that you needed bazillions of memory for good results. I’m very out of the loop.

[-] Greg@lemmy.ca 3 points 1 day ago

I have a RTX4090 (24GB VRAM) plus 128GB DDR4 system RAM on Ubuntu 24.04 and I am running these models locally

| Model | Speed | Context window | |


|


|


| | Qwen3.8-Flash-Next UD-Q4_K_XL — 125B/6B-active hybrid MoE, 104 GB | ~16.7 tok/s | 128k tokens | | GLM-5.3-Flash UD-IQ3_XXS — 320B/18B-active MoE, 120 GB | ~9 tok/s | 128k tokens | | Qwen3.8-27B UD-Q4_K_M — dense hybrid, 16 GB, fully on the 4090 | ~49 tok/s | 262k tokens |

The larger models don't fit on the GPU alone but they're mixture-of-experts and their active weights easily fit on the 24GB. The bottleneck is my system RAM as llama.cpp has to constantly move the experts between RAM and VRAM. I'm running older hardware, AM4 CPU, DDR4 RAM, etc. so 128GB is my limit on system memory and it's relatively slow compared to DDR5. With DDR5 I would expect a good bump in speed for Qwen 3.8 Flash Next and GLM 5.3 Flash.

I also have a Kimi subscription which I use as my daily agentic driver and use DeepSeek for random tasks (DeepSeek Flash 4.1 is really fast and cheap).

[-] gorbinos_quest@lemmy.world 2 points 1 day ago

That’s great. I’m on a 4090, but only 64GB of system RAM (DDR5 though). I might give it a go.

[-] WolfLink@sh.itjust.works 1 points 1 day ago

Interesting. I didn’t know how that worked, but I might give it a try. I have a 3090 with a lot of RAM.

[-] chrisashtear@lemmy.zip 3 points 1 day ago

Yes, pretty well. Qwen 3.8 can be run on one 3090 at Q4. Q8 runs on my computer with 48g of vram. They can do 128k to 256k context

[-] vapeloki@lemmy.world 2 points 1 day ago

It depends. If you have enough host memory, those models are pretty efficient and work well with MTP.

AMD Strix is good for it. Slower memory but a lot of it, no CPU Fallback.

And the results are not worse then Claude at least

[-] gorbinos_quest@lemmy.world 1 points 1 day ago

That’s nice, I might give it a go. I thought Claude would be way better than those models.

[-] vapeloki@lemmy.world 2 points 1 day ago

Standing rule in my projects:

  • never read code from imported libraries/modules
  • read all of our design and code style rules before even responding to a question.

Claude never followed those rules. Or the design guides or the style guides.

Qwen adheres to every fucking single word.

[-] vapeloki@lemmy.world 4 points 1 day ago

We have a big pitch coming. And I am happy that I could convince our uppers to try a stund.

We have a hackathon with a customer. One of our teams will work on strix workstations. We local, custom tuned models, run directly from a small solar installation.

We will reveal this at the end. Our power consumption and the 0bits into USA metric

[-] Greg@lemmy.ca 2 points 1 day ago

Our power consumption and the 0bits into USA metric

That's amazing! Keep us in the loop about the results

this post was submitted on 14 Sep 2026
58 points (98.3% liked)

Fuck AI

8217 readers
1759 users here now

"We did it, Patrick! We made a technological breakthrough!"

A place for all those who loathe AI to discuss things, post articles, and ridicule the AI hype. Proud supporter of working people. And proud booer of SXSW 2024.

AI, in this case, refers to LLMs, GPT technology, and anything listed as "AI" meant to increase market valuations.

founded 2 years ago
MODERATORS