this post was submitted on 13 Aug 2026
201 points (98.1% liked)

Programming

28091 readers
317 users here now

Welcome to the main community in programming.dev! Feel free to post anything relating to programming here!

Cross posting is strongly encouraged in the instance. If you feel your post or another person's post makes sense in another community cross post into it.

Hope you enjoy the instance!

Rules

Rules

  • Follow the programming.dev instance rules
  • Keep content related to programming in some way
  • If you're posting long videos try to add in some form of tldr for those who don't want to watch videos

Wormhole

Follow the wormhole through a path of communities !webdev@programming.dev



founded 3 years ago
MODERATORS
 

studies show a clear trend – output is up (more code, more commits, bigger diffs), but outcomes don’t reflect that trend. If anything, the average team is taking longer to ship worse software

top 50 comments
sorted by: hot top controversial new old
[–] Simulation6@sopuli.xyz 8 points 13 hours ago

LLMs can't tell the difference between code and comments in some cases. Older code bases that have had a number of hands touching it over the years have a lot of commented out code and dead methods. LLM sees that as good live code. Also there will be comments like 'this is a stupid way to do this, but I don't have time to fix it right now'. LLM has no sense of humor in this case.
And sure. I should clean all this up before hand, but who has time for that? Project has one, part time developer working on it now.

[–] NigelFrobisher@aussie.zone 15 points 1 day ago (1 children)

I could’ve told them this, tbh.

[–] douglasg14b@lemmy.world 8 points 19 hours ago (2 children)

Anyone can tell anyone anything, it doesn't make it true or validated.

[–] padreug@programming.dev 1 points 11 hours ago

don't forget, even a broken clock is right twice a day 😄

[–] Feyd@programming.dev 104 points 1 day ago (2 children)

output is up (more code, more commits, bigger diffs)

We've known measuring output by lines of code is counterproductive for a long time.

[–] jtrek@startrek.website 41 points 1 day ago (2 children)

Has management known that?

A lot of problems seem to be downstream from "management are idiots and assholes"

[–] dreamkeeper@literature.cafe 3 points 11 hours ago

At any minimally competent company they are aware.

However my company mostly has former engineers as engineering managers.

[–] Feyd@programming.dev 8 points 1 day ago (3 children)

I'm sure there were pockets but it legitimately seemed like that dragon had been slain until recently. Trying to assign more meaning to scrum points has been in vogue for a while though.

[–] jtrek@startrek.website 6 points 1 day ago (1 children)

My team assigns both hours and points to tasks. I've never seen the points used for anything but they still spend time on it.

[–] Kissaki@programming.dev 3 points 17 hours ago

That's... an interesting approach.

That battle is perpetual. Lazy management sees a number that resembles a statistic, and try to use it as an easy metric for stuff it doesn't represent. Points do aggregate into velocity, which is worth measuring. But on their own, points are a proxy for estimation in $SPRINT_LENGTH days. The way to manage up is to keep making it clear that the smallest unit of estimation in Agile is the sprint length; points are used to subdivide that but only to ensure that the sprint itself is not overloaded and thus an accurate estimate.

load more comments (1 replies)
[–] criss_cross@lemmy.world 13 points 1 day ago (1 children)

Can you retell my company that?

[–] Kissaki@programming.dev 5 points 17 hours ago

For a small exorbitant consulting fee that legitimizes me, I can tell them that.

[–] chicken@lemmy.dbzer0.com 17 points 1 day ago (1 children)

LLMs cannot distinguish between recent and out-of-date information in the context, and information in the model itself, learned during training (“dominant priors”), can often “outweigh” information we give it

LLM inference is more accurate when we give them examples (demonstrations) rather than just describing what we want.

Deep neural networks, including LLMs, struggle to learn patterns with long-range dependencies, at any scale of model. They will always be “driving in fog”, with local, short-range probabilities crowding out long-range ones. In case you were wondering why they suck at the “big picture” – probabilistically, it’s a blur.

I'd guess that what all this stuff adds up to is, sustainable use of LLMs as a coding tool for nontrivial projects calls for an entirely reworked set of software development practices to conform to its limitations effectively, but the people in charge really really want and believe it to be a drop-in efficiency boost, and a big mess results. This reminds me a lot of the articles and arguments I've read over the years about low level vs high level programming languages and frameworks. Probably will play out a similar way.

[–] douglasg14b@lemmy.world 2 points 14 hours ago* (last edited 14 hours ago)

A lot of tokens tends to help move the needle, building in guardrails, sophisticated review of both in process work & final outputs, and consensus helps raise the bar considerably.

That and sane, universally consistent, well structured, low tech debt, organizationally elegant, systematically evolved codebases. Something the grand majority of teams don't have to begin with.

Nevermind being well documented, with excellent opinionated linting and static type enforcement configurations, with robust CI checks. More things many projects don't seem to have.

Turns out that giving AI poorly engineered and maintained codebases just amplifies the poor engineering and low rigor already present. Something that's endemic to our industry.

[–] ArseAssassin@sopuli.xyz 74 points 1 day ago (2 children)

One study found a significant correlation between confidence in AI output and belief in the paranormal.

💀 💀 💀

[–] 87Six@lemmy.zip 11 points 1 day ago

That's HILLARIOUS

[–] MonkeMischief@lemmy.today 3 points 1 day ago* (last edited 1 day ago) (1 children)

Lol that's funny. My belief in having a divinely created soul is exactly why I think humans can't be replaced by these supercharged drunken parrots. :O

...But yeah this is probably referring to the ones who, under the right conditions, will start to sense "ghosts in the machine" and go down the rabbit hole to chatbot psychosis...

[–] Senal@programming.dev 8 points 1 day ago* (last edited 1 day ago) (1 children)

I am 100% genuinely not trying to start a fight, this a legitimate question to which I really would like to know the answer.

What is the difference between a divine creation and the "ghost in the machine "?

I will give context , faith is genuinely confusing to me.

I have no issue with personal faiths unless you are trying to force everyone to have the same faith as you.

I consider anything above small scale organisation of religion to be the worst thing that can happen to faith, because people are people and power corrupts.

I suppose my real question is , for a concept that prohibits proof as part of it's definition, how do you determine that one faith (or system) is better than another?

[–] arendjr@programming.dev 7 points 1 day ago (1 children)

I’m not the person you responded to, but I can try to take a stab at that…

for a concept that prohibits proof as part of it's definition, how do you determine that one faith (or system) is better than another?

Objectively, you don’t. But people’s experiences go beyond the objective, and we all have our subjective experiences too. Faith is how we make sense of those.

The question then isn’t, which one is better, but whose other experiences and faiths do we relate to? We build friendships and alliances based on those. Not because they’re better, per se, but because we feel comfortable or safe with them.

From that perspective it makes a lot of sense that we value other humans, because they can make us feel understood and appreciated, and we can exchange ideas with them which in turn refines our faith. Now, some people get tricked into thinking you can do the same with machines, but well, I think you can guess where I too stand on that idea.

[–] MonkeMischief@lemmy.today 4 points 1 day ago

That was a very insightful answer. Well said! Thank you very much for replying. :)

I will have to contemplate a little bit, and respond to the question myself as well.

[–] dejected_warp_core@lemmy.world 28 points 1 day ago* (last edited 1 day ago) (1 children)

There are some cardinal sins within the annals of software development in the workplace. The relevant two here are:

  • Do not build a scoreboard for productivity
  • Do not equate keystrokes with effort or value rendered

These create perverse incentives that can really screw everything up, including creating permanent damage to company culture, products, and productivity. Anytime you see these things done, it's because you have (or are) lazy-ass management.

Edit: I also just learned that this is an application of Goodhart's Law.

[–] themaninblack@lemmy.world 5 points 16 hours ago* (last edited 16 hours ago)

Oh god a former company had the burndown chart on a big plasma screen at all times

[–] ell1e@leminal.space 23 points 1 day ago* (last edited 1 day ago) (2 children)

This doesn't seem to cover there is also no LLM that doesn't plagiarize, or where the training data appears to be compatible with such behavior (e.g. CC0). Now I don't know what that means legally, but morally it seems to be tossing away other project's licensing and I think for FOSS as a whole that's no good.

Also something worth reiterating: https://machinelearning.apple.com/research/illusion-of-thinking LLMs apparently can't do basic logical reasoning. Even a junior coder can do that. I'm always surprised anybody would let LLMs near their code, at all.

[–] blarghly@lemmy.world 5 points 1 day ago (1 children)

Of course it didnt cover that. This was an article on the efficacy of AI in writing software, not a treatise on ethical or legal concerns. I would completely lose trust in the author if they started piling on every reason why "AI Bad", because it would be clear that they have an agenda

[–] ell1e@leminal.space 1 points 10 hours ago

In my opinion that's not the framing the article uses. It just says "I’m currently pulling together a bunch of sources – that are mostly recent – on the topic of LLMs and their use in software development.".

Leaving out any ethical concerns for that basic framing seems like a pretty notable choice, so I thought it was worth pointing that out.

I recommed you reading this

It summarizes really good not only the moral, but also the legal problems of AI, vibecoding and "AI-assisted/AI-boosted" programming/engineering/development.

[–] hoshikarakitaridia@lemmy.world 27 points 1 day ago (1 children)
[–] WhatAmLemmy@lemmy.world 21 points 1 day ago* (last edited 1 day ago) (2 children)

It's obvious if you actually use software beyond the average literacy of a talking chimp. I use hundreds of apps across iphone, mac, and linux os's. Literally none of them have noticeably increased in quality, stability, or feature-set beyond their average between 1-5 years ago.

Mac and iphone appear to have more bugs and shittier quality control than at any other point in the last decade.

Quality software is getting harder and harder to find thanks to all the slop-abandonware being promoted by slop-content and slop-SEO on slop-enshittified search engines.

I notice far more idiocracy-grade errors in digital content, cx, business processes, product listings, etc than ever before.

Weather forecasts have gone to dogshit in the last 2 years. Even same-day forecasts can shift on a dime unpredictably. It's at the point where I check 3 apps. Until this year, I never had a time where I woke up to 0% chance of rain and sunny, then looked outside to see rain. Not a sun shower. A rainy day hour-long downpour.

Auto-generated subtitles are great for content that was never going to receive human attention, but they're clearly being used to replace humans. At least once a week I notice a major contextual error that completely alters the perception of the line/scene, and there's no way to submit corrections.

Art, culture, and knowledge are being actively corrupted, bastardized, and destroyed.

The future simultaneously sucks while being dumb as all fuck. Complete clown show run by the idiot criminals, rapists, pedophiles, scammers, and thieves.

[–] eah@programming.dev 5 points 1 day ago* (last edited 1 day ago) (4 children)

Weather forecasts have gone to dogshit in the last 2 years. Even same-day forecasts can shift on a dime unpredictably. It’s at the point where I check 3 apps. Until this year, I never had a time where I woke up to 0% chance of rain and sunny, then looked outside to see rain. Not a sun shower. A rainy day hour-long downpour.

Supposing you're right that weather forecasts have gone to dogshit (I'm skeptical because your evidence is anecdotal), there may be reasons beside AI for that.

[–] WhatAmLemmy@lemmy.world 3 points 21 hours ago* (last edited 20 hours ago)

Yeah I was gonna add the weather part could be entirely explained by climate change and fascisms war on science, but forgot.

I'd never heard of the 5G impact, so thanks for that. I consider it unlikely unless there were a significant change in 5G spectrum usage during the last year. 5G has covered most Australian major cities since 2021, and 5g coverage is essentially non-existent outside of populated areas (90% or more of Australia).

Weather forecasting is actually one area where AI should far exceed what humans could ever possibly achieve without it. It is impossible for humans to process and adapt to the changing patterns across hundreds or thousands of variables interacting with each other in real time.

load more comments (2 replies)
[–] Feyd@programming.dev 7 points 1 day ago (1 children)

Yes it blows my mind everyone can't see all the historically perfectly fine software starting to crumble to dust

[–] The_Decryptor@aussie.zone 1 points 22 hours ago

The big one for me was rsync, maintainer started vibecoding it and the first release with those changes had a huge amount of regressions.

Yet it keeps getting held up as an example of "using AI tools properly".

[–] MagicShel@lemmy.zip 16 points 1 day ago (3 children)

This aligns with my experience, largely. Of course it's still my job to maximize LLM effectiveness within my organization. Which is a delicate balancing act to protect my teams from overeager executive leadership looking for huge gains.

My own summary is that AI can be an accelerator, but the harder you lean into it, the worse outcomes will be. No matter how much code is written, you still need actual human minds to understand it and they can only handle so much volume before getting overwhelmed.

Also, if AI gives you 20% productivity gains, but that 20% goes into playing with AI trying to get more, you haven't really gained anything. Usage needs to be standardized rather than developers constantly negotiating with AI trying to coax out better outcomes.

[–] blarghly@lemmy.world 7 points 1 day ago

The impression I've gotten, fooling around with it at home and talking to friends in tech and hearing from actual users online is:

A good developer can develop faster with it. Giving it small, discrete tasks for first drafts or throw away code (like bash scripting) can work well.

It's better google. If you are trying to figure out if a function that will do a thing exists, or are trying to figure out what architecture would work best in a given situation, it can be helpful. But in these cases, it should be used carefully - dont ask it to do your work for you, ask it to give you options, pros and cons, and sources. But in this regard, it can do a lot to help an experienced developer become more productive faster in a stack or tool they are unfamiliar with.

Vibe coding is a real thing, and it can work. For internal tools in a small company, a non-technical person can create a mostly functional piece of software to get a job done. My expectation is that over time, these people will become real developers, as they end up dealing with bugs and edge cases in the vibe code they created.

At the top end of ai-for-software-development, there is some sort of something with automated iterative looping and verification, where a developer can translate a set of requirements into code, and then a collection of ai agents iteratively develop the code until it works as expected. This is what the tech bros seem really hyped on, and it does seem to work... but at the same time, my feeling is that this is how you get multimillion dollar AI bills. And presuming this is how big tech is developing their products - it seems prone to making inefficient, buggy code, so I don't think it will be worth it long term.

[–] jerakor@startrek.website 3 points 1 day ago

This is a tale older than AI. Most of the AI productivity pushes I struggle to get adopted fail not because of AI bad or its too hard to do. They fail because of a broken CI/CD pipeline. They fail because some team thinks their process is sacred and unique.

load more comments (1 replies)
[–] JeeBaiChow@lemmy.world 10 points 1 day ago

Lol. Who knew?

[–] LemmyBruceLeeMarvin@lemmy.ml 2 points 1 day ago

Someone's ITIL certified

[–] faltryka@lemmy.world 8 points 1 day ago (4 children)

Some of this does not line up with my lived experience pretty starkly.

Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.

Not perfectly well, but llms are designed specifically NOT to be perfect deterministic executioners. Still though, pretty well.

I have seen that in a jr engineers hands llms get to bad outcomes fast, and unintuitively (to leaders…) usage of llms in coding does not provide a path for a he engineer to upskill into a sr engineer. A sr engineer with llms though is almost always radically augmented regarding their output speed on task completion.

[–] MagicShel@lemmy.zip 10 points 1 day ago

I agree with your last paragraph. We had about 6 weeks of unlimited AI spend before the costs reached executive leadership, and in that time I saw the least experienced developers spend the most with the least to show for it.

But I will say that another factor is thinking that if you get 10% gains from a little AI, then a lot of AI will get you 100%.

But I find the article is right about repo-wide docs. At least on their own. I find having small markdowns (often in the form of skills/commands), focused on specific tasks reduces spend (especially when your execution agent is a low cost model, leaving the reasoning to dedicated agents) and gives better outcomes. Loading massive docs into every task reduces the attention to the task at hand and often confuses AI as the reasoning part of the model becomes overwhelmed and starts inferring wrong things confidently.

I suppose it heavily depends on the scale of the repo though. A large microservice with multiple upstream services it needs to call spends a lot tokens on API which is unnecessary for most tasks. And then it decides to use the wrong one.... I have stories lol.

[–] rimu@piefed.social 2 points 1 day ago

It's possible to win lots of battles but still lose the war. You can ask Trump about that :)

[–] melfie@lemmy.zip 1 points 23 hours ago (1 children)

Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.

Same. AGENTS.md files and the like are quite effective. Especially if you’re reviewing the code and making the LLM help you update the markdown files when it makes a mistake to prevent the same type of mistake in the future. Having concrete examples of “good” vs. “bad” to illustrate each architectural rule goes a long way.

For any feature or bug fix that is “painting with the colors already in the tray”, it makes sense to let a LLM write the code. Humans will introduce new tech and new patterns out of boredom and turn the codebase into a big Frankenstein, but the LLM will just follow the architectural guidelines indefinitely.

[–] faltryka@lemmy.world 2 points 22 hours ago

Agree, I have them curated lessons.md anytime they make a mistake and have found that to be highly effective. Every now and then a lesson goes defunct and needs pruned, but I think that’s just part of the new swe skill set.

load more comments (1 replies)
load more comments
view more: next ›