Puzzled

I’ve encountered a number of pieces in which the authors wrote that LLM AI never disagreed with them. That puzzles me.

I occasionally use ChatGPT, Claude, Grok, and Gemini. They’ve all disagreed with me at one time or another. When I have remonstrated with them they occasionally back down, particularly when I point out the logical flaws in their statements. Even more occasionally they’ve been flatly wrong. When I point that out, they’ve acknowledged that.

One of the great things about contradicting an LLM AI model is that I can do so harshly without worrying about its feelings.

I don’t think the proper lesson from this is that the models aren’t useful. It’s that they are easy to misuse.

8 comments… add one
  • walt moffett Link

    Wonder if they be using AI as an amen section.

  • I certainly am not. I use it as a proofreader, editor, and artist. I also use it to toughen my arguments.

  • steve Link

    I use it as a quick check for data and then do my own double checking. Also, as I am not a particularly good writer, Im sure you have noticed, if I am doing something important I ask it for a rewrite. I think the number of errors is decreasing. I really dont ask an LLM for its opinions on topics.

    OT- I dont think you have ever written on the fertility issue, but you do touch on Russian issues. The following quote is from a Cowen interview with Ioffe.

    “Six years before the full-scale invasion of Ukraine, Russia decriminalized domestic violence. When I was reporting on this for the book, I just couldn’t wrap my head around it. I kept asking lawyers, activists, victims, or survivors, “What was the point of this? Why decriminalize domestic violence?” Some said, “Oh. It’s to reinforce the traditional family order that the man is the head of the household, that he’s allowed anything, and that this is the traditionalist way that husbands behave. They beat their children, they beat their wives, and that’s how they enforced order in the home.”

    Of note, in spite of increasing payments for having a kid and a big effort to promote fertility, the TFR continued to drop. I am not totally surprised Russian men would want this. I had a number of problems with eastern European male staff I hired, nearly all interpersonal stuff. My boss asked me to stop hiring them.

    Steve

  • walt moffett Link

    Not suggesting you are, just the folks who say AI never disagrees, they forget the advice about flatterers. My own use of AI is mainly through DuckDuckGo’s search summaries which give the source. everything else, is a happy slog wearing a snorkel and hip boots.

  • I have a strong opinion on the fertility issue but it is so contentious I don’t bother airing it.

  • steve Link

    More contentious than the Russians who seem to think alternately paying the women and beating them will get more kids? Hard to believe.

  • PD Shaw Link

    Given how frequently AI LLM’s are wrong, I think Dave’s experience just suggests that these tools cann’t be used by anyone that doesn’t already have a pre-existing knowledge base to cross-examine.

    A couple of weeks ago a commentor on another blog claimed that Lincoln’s selection of Johnson for VP was his greatest mistake. Maybe? Did it happen? No. I asked Gemini if Lincoln chose Johnson as his running mate, and it answered “Yes, Abraham Lincoln chose Andrew Johnson to replace Hannibal Hamlin as vice president.” The source appears to be Reddit, Quora and some superficially relevant links. I tried a few different queries but since I knew one of the best Lincoln historians had downloaded extended chapters of his Lincoln bio on a university website, I asked “what does [author] think of Lincoln’s role in the selection of Andrew Johnson,” and I was told that the author shares the consensus among leading historians that Lincoln did not actively direct or orchestrate the choice, linking to a pdf chapter of the author’s book, which concluded that “Johnson turned out to be a disastrous choice, but Lincoln had nothing to do with his selection.” Re-running the query this morning, Gemini still gives the wrong answer, saying that Quora discussions are where users share a consensus on the topic.

    I imagine a few things are happening, such as an assumption that Vice Presidents were selected in the same way in the past as in the present, American political histories attribute far too much to the President as the primary if not sole force in politics, or most discussions are really more interested in why Johnson was deemed a good choice and give minimal attention to the process. But mainly the issue presents negative space that Gemini has filled with meager material.

    I don’t know if this is AI sycophancy, but it suggests that perhaps it understood a common perception has a life of its own. Probably a lot of high school history books state or assume Lincoln made the selections. Maybe it assumes the wrong answer anticipating that the real question (why?) would be the next prompt. But the dynamic is enforcing common error, making it more common.

    If its wrong about things you know, how can it be relied upon to help understand subjects of curiosity? The best cross examination, doesn’t merely point out that the witness was wrong, but leaves the jury believing the witness is not credible, and perhaps opposing counsel is not reliable as well.

  • Dave’s experience just suggests that these tools cann’t be used by anyone

    The irony is that many of those who might actually use AI effectively have no inclination to do so. Every junior developer is starting to claim that he or she is an AI expert. The number of people who could use AI effectively and have the experience and knowledge to use it is not only small but managers are mistrustful of them.

Leave a Comment