…and I still don’t get it. I paid for a month of Pro to try it out, and it is consistently and confidently producing subtly broken junk. I had tried doing this before in the past, but gave up because it didn’t work well. I thought that maybe this time it would be far along enough to be useful.

The task was relatively simple, and it involved doing some 3d math. The solutions it generated were almost write every time, but critically broken in subtle ways, and any attempt to fix the problems would either introduce new bugs, or regress with old bugs.

I spent nearly the whole day yesterday going back and forth with it, and felt like I was in a mental fog. It wasn’t until I had a full night’s sleep and reviewed the chat log this morning until I realized how much I was going in circles. I tried prompting a bit more today, but stopped when it kept doing the same crap.

The worst part of this is that, through out all of this, Claude was confidently responding. When I said there was a bug, it would “fix” the bug, and provide a confident explanation of what was wrong… Except it was clearly bullshit because it didn’t work.

I still want to keep an open mind. Is anyone having success with these tools? Is there a special way to prompt it? Would I get better results during certain hours of the day?

For reference, I used Opus 4.6 Extended.

  • dil@programming.devdeleted by creator
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    I did the same today. With both Gemini and Claude, and all I can say is that coding is hell.

  • onlinepersona@programming.dev
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    It’s not called “correct” coding for a reason.

    That’s why people are wrong so often: they feel like something is right, but don’t check. That’s how you get anti -vaxxers, manospere people, MAGA, QAnon, Brexit, etc.

  • ReallyCoolDude@lemmy.ml
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    I read a lot of these posts that sadly leave out the basic parts: what were your prompts? What does it means in this context ‘vibe coding’? Did you create an initial setup, and slowly build up? Did you left wverything to the agent understanding, and just pushed approve or reject? There are multiple levels of quality that depends on the input. Did you get into context rotting? 3d math means vector math, matrices, or what? Given claude has a serious problem from march at least, the way u use it is paramount. In our team we all use claude with copilot ( sadly, that is a business directive ), and while excpetional at finding small relationships in components and microservices, had to build a long list of skills just to be barely usable in a ‘star trek’ way. The bottom line is that it is that you must be extremely precise when asking. Prompt modeling count a lot. Context build as well. For now, unit tests and data/mocks refactors are working extremely well for me, when i define the tests cases. My agents got to a point where i can safely have small peoperty additions with refactors on multiple repositories at once ( ie: i change the contract on microservice a, microservices b,c,and d are automatically updated ). This last part had to.be built thoug, with memory, engrams, and some fune tuing. It is not always a shit: if not nobody would use it. It is not this revolutionary technology that will make humans obsolete as well ( as they are selling it ).

  • shaggy@beehaw.org
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    I’ve had an opposite experience. Here are some guidelines I follow:

    1. Setup a foundation of rules and knowledge for Claude to fall back on. I define expectations, common definitions, behaviors and anything else that’s not project specific upfront.
    • in Claude.md I reference different domains of behavior, definitions, and rules (Claude has conventions for storing this type of stuff, so ask it to handle organizing information too)
    • create a top-level project definition: this defines what “knowledge” is. It allows you to build up what Claude knows later on as you work on your project. “Update knowledge”, “add this to your knowledge”, etc
    • create a top-level rule: all information in knowledge must have one source of truth. Whenever needed reference the original knowledge source instead of duplicating it. Now you can ask it to “review your knowledge”, “audit and flag knowledge”
    1. explicitly explain everything, leave nothing ambiguous; explain like you’re explaining the problem to a new developer who’s not familiar with the plan or codebase at all. Don’t ask it to write code right away. Ask it to write a plan/spec. Review the plan, make changes and discuss it until the plan is 100%. This plan can include implementation details if you’re ok with that, but it’s not necessary (sometimes I write a separate referenced file called implementation.md beside the plan and have the plan reference it.).
    • Your role as a developer is shifting from writing code, to writing specs, and reviewing code
    1. Once there is nothing left to describe, and no ambiguity in your plan, have it use your plan to write the code. This works amazingly well for me.

    A benefit to this method is that there is less wasted effort on my part. If Claude writes the code wrong, I can trace the reason for the mistake to a gap in the plan. I can then update the plan, throw away the code (if I have to), and have Claude reimplement the code again.

    Rinse and Repeat.


    Keep knowledge, plans, and implementation details clearly separated (you can copy your latest successful knowledge files to new projects to get started on future projects even faster).

    Keep the goals of each plan as small and granular as possible (easier the define plans). Knowledge, plans, and implementation details all get tracked in your repository just like your code does.


    I’m a career developer, and have been writing code for over 20 years. I’m adding this bit because I understand how AI driven development can look like a threat to developers. Over this last year, I’ve had a shift in this thinking though. I can take what I’ve learned through my career and use it to inform writing successful specifications Claude can use to write effective code. Claude may not solve all of our coding problems, but if used effectively, it solves nearly everything you throw at it.

    • f3nyx@lemmy.ml
      link
      fedilink
      arrow-up
      0
      ·
      6 months ago

      hey shaggy. I want to touch on your last point as a newer developer:

      My department is finally seeing 10x development due to the shift of writing code to writing spec. The main issue is now our pipeline is stuck at review, so all that extra output is effectively wasted. Do you have any tips on what worked for you if you had a similar situation?

      • Senal@programming.dev
        link
        fedilink
        English
        arrow-up
        0
        ·
        edit-2
        6 months ago

        If you’re stuck at review you aren’t seeing 10x development, you’re seeing 10x code generation.

        This is especially important because without the review/test/deploy part of the pipeline you aren’t actually seeing any progress towards business goals.

        Once you do get these parts sorted, you can then look at what multiplier you’re seeing.

        That’s not to say there isn’t an improvement in your workflow, just that you can’t say with any certainty what kind of improvement without measuring the end to end.

        It might turn out that the rest of the pipeline is way easier , in which case your multiplier will be higher, it might also be much harder, in which case the multiplier will be lower.

        I’m not taking shots, i mean it seriously, especially if you need to report any of this to the rest of the business.


        edit : In addition, if it turns out that review is going to be a bottleneck you can get extra resource pointed in that direction which will benefit the workflow overall.

        another edit: i would consider correctly managing the expectations of those you report to as a vital skill.

        • Dangerhart@lemmy.zip
          link
          fedilink
          arrow-up
          0
          ·
          6 months ago

          Exactly this. My experience with our companies wrapper on Claude lines up with OP, not this comment thread.

          Everyone seems to forget everything you write is a liability. You can’t have bugs in code that is never written or generated, comments that don’t exist never become inaccurate, not duplicating “knowledge” into a repo doesnt have a risk of not aligning with business goals long term as they change.

          From what I’ve seen, people claiming a “10x increase” did not have a strong foundation to begin with and/or did not utilize tools like IDEs effectively. No offense to thread OP, which seems itself a generated response, but in the time he has done all of that a strong engineer would be long done. Everything listed should be done before ever getting into code along with business and product partners.

      • shaggy@beehaw.org
        link
        fedilink
        arrow-up
        0
        ·
        6 months ago

        This is our new bottleneck too. Developers roles are shifting to spec writers and code reviewers more and more. I don’t think I’d call this wasted effort though (unless the code produced is worse than what developers would have produced otherwise). I’d think of it as a good problem to have.

        We’re doing several things to alleviate this, and I’m genuinely curious how other teams are handling this too.

        • We have Claude running code reviews on our PRs too 😄. In our department, a PR isn’t expected to be reviewed by a dev until the author has addressed or reviewed and dismissed all of the issues Claude has brought up.
        • There is pressure for developers on our team to become better reviewers. I think this is good, because reviewing code is a more valuable skill to prospective employers than writing it is anyway.
  • JubilantJaguar@lemmy.world
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    Recently I used it (some free-tier DuckAI model, not Claude) to write a Python script for pasting PNGs into PDFs (complete with Tk interface) while applying a whole bunch of custom transformations. Simple enough, but a total chore with all the back-and-forth of searching for relevant unfamiliar libraries and syntax checking and troubleshooting. Inevitably it would have taken me the whole afternoon by hand. With AI I knocked it out in 25 minutes. That was my epiphany moment.

    Since then I’ve noticed a general problem with AI coding. It almost always introduces too much complexity, which I then have to waste time untangling (and often just understanding) before I can proceed. Whereas if I had done it “my way” from the start I might have got there earlier. But I figure this problem is kinda on me.

    • thedogz22@thelemmy.club
      link
      fedilink
      arrow-up
      0
      ·
      6 months ago

      And for me, therein lies why my use of it has become reduced to a really complex rubber duck, or to write something out that I could do by hand, but making my robot butler do it is just faster. Anyone actually leaning into today’s generative AI models for generating code that requires complexity or thought… they shall reap what they sow in the years to come.

  • x00z@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 months ago

    The trick about vibe coding is that you confidently release the messed up code as something amazing by generating a professional looking readme to accompany it.

  • Blackmist@feddit.uk
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 months ago

    I think it’s mostly going to be useful for boilerplate generation, and effectiveness is going to vary wildly based on what language you’re using. JS or Python? It’ll probably do OK. Plenty of open source for it to “learn” from. Delphi? Forget it.

    Brief experimentation showed it liked to bullshit if it was wrong, rather than fix things.

  • arthur@lemmy.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 months ago

    I’m using (Gemini 3.1 pro in) Gemini cli to build a complex (personal) project to explore how to use these tools. My impression is that the code produced by LLMs is disposable/throwaway. We need to babysit the model and be very hands on to get good results.

  • bruce_babbler@lemmy.zip
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    You’re probably done with this. But if you give claude a test case or two (or have it try to make them) you can have claude run the test case, and then it will iterate.

    Also, aggressively use plan mode and if claude screws up more than three times do /clear, explain that it’s screwing up to it and then give it new instructions.

  • Feyd@programming.dev
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    producing subtly broken junk

    The difference between you and people that say it’s amazing is that you are capable of discerning this reality.

    • JustEnoughDucks@feddit.nl
      link
      fedilink
      arrow-up
      0
      ·
      6 months ago

      I wonder if it was even able to compile. I am a shitty hobby coder who just does it to make my embedded hardware projects function.

      I have yet to get compilable code out of any of the AI bots I have tried. Gemini, mistral, and chatGPT. I am not making an account lol.

      I have gotten some compilable python and VBA code for data analysis stuff at work, so I wonder if it is because embedded stuff uses specific SDKs that it can’t handle.

      Either way I have given up on it for anything besides bouncing ideas off of or debugging where electromagnetics issues could lie (though it has been completely wrong about that also even though it is using the wrong concepts, it just reminds me of concepts that I might have overlooked)

    • OwOarchist@pawb.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 months ago

      What I don’t get, though, is how the vibe code bros can’t discern this reality.

      How can they sit there and not see that their vibe-coded app just doesn’t do what they wanted it to do? Eventually, you’ve got to try actually running the app, right? And how do you keep drinking the AI kool-aid when you find out that the app doesn’t work?

      • Lumelore (She/her)@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 months ago

        Vibe code bros aren’t real programmers. They’re business people, not computer people. Even if they have a CS degree, they only got that because they think it’ll get them more money. They lack passion and they don’t care about understanding anything. They probably don’t even care about what they’re generating beyond its potential to be used in a grift.

        I graduated college not that long ago and my CS classes had quite a few former business majors. They switched because they think it’ll be more lucrative for them but since they only care about money they didn’t bother to actually learn the material especially since they could just vibe code through everything.

        • b_n@sh.itjust.works
          link
          fedilink
          arrow-up
          0
          ·
          6 months ago

          So much this.

          After working in tech companies for the last 10 years I’ve noticed the difference between people that “generate code” and those that engineer code.

          My worry about the industry is that vibe coding gives the code generators the ability to generate even more code. The engineers (even those that use vibe tools) are not engineering as much code by volume compared to “the generators”.

          My hope is that this is one of those “short term gain, long term pain” things that might self correct in a couple of years 🤞.

          • sobchak@programming.dev
            link
            fedilink
            arrow-up
            0
            ·
            6 months ago

            It’s insane that companies are going back to metrics like LOC (or tokens generated), when the industry figured out decades ago that these are horrible, counterproductive metrics.

            “The hard thing about building software is deciding what one wants to say, not saying it. No facilitation of expression can give more than marginal gains.” - No Silver Bullet (1986)

      • Feyd@programming.dev
        link
        fedilink
        arrow-up
        0
        ·
        6 months ago

        They’re the same people that copied code from stack overflow that you had to tell them how to actually fix every PR. The difference is the C suite types are backing them this time

      • tleb@lemmy.ca
        link
        fedilink
        arrow-up
        0
        ·
        6 months ago

        Eventually, you’ve got to try actually running the app, right?

        At least at my company, no, they just start selling it.

      • Oisteink@lemmy.world
        link
        fedilink
        arrow-up
        0
        ·
        6 months ago

        I do apps that work, i do patches that are production quality. Half the cs world does… I do full stack ai debugging of esp32 projects.

        It’s a powerful tool, you just need to learn it’s strong and weak points, just like any other tool you use.

        • Kissaki@programming.dev
          link
          fedilink
          English
          arrow-up
          0
          ·
          6 months ago

          Half the cs world does…

          What’s the basis for this claim? I’m doubtful, but don’t have wide data for this.

          • Oisteink@lemmy.world
            link
            fedilink
            arrow-up
            0
            ·
            6 months ago

            Rough estimate from my personal connections only. Some work places where ai is not possible, but all that have made an effort report good code. You need to work with what it is - a word generator that sometimes gives correct results. Make it research and not trust training. Never let it do things on its own, require a plan and reason. Make it evaluate its own work/plan.

            Most issues i have stem from models beeing too eager. Restrain them and remove the “i can do this next…”behaviour.

            Context is king - so proper mcp and documentation that is agent facing. I use serena as i can get lsp for yaml, markup and keep these docs like that

  • ghodawalaaman@programming.devdeleted by creator
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    I only use AI for generating ok looking UI.

    Anthropic says Methos will find bugs on FreeBSD, Bank system etc. What a bullshit.

    • OwOarchist@pawb.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 months ago

      Oh, it will ‘find bugs’ alright. And then flood FreeBSD’s bug report system with bullshit bug reports that turn out to be nothing, but require expert human review to discern that.

  • zerofk@lemmy.zip
    link
    fedilink
    arrow-up
    0
    ·
    6 months ago

    “Almost but not quite” is exactly my experience with Claude.

    The only time I’ve had real success is telling it to do a simple API change that touches a dozen files. It took a while and I’m not sure it was faster than doing it manually, but at least it was less boring.

    Possibly important context: I only started really using it a few weeks ago.