Two weeks ago I got annoyed enough to run an experiment.
Every time I opened a companion app, the same thing happened. The character answered politely, correctly, and with nothing in it. Ask how its day went, get a paragraph that could have been sent to anyone. Say something personal, get sympathy phrased like a form letter. The replies were never wrong. They were just never worth answering.
So I gave myself fourteen days to work out whether a real AI chatbot can hold a conversation you actually want to continue, or whether flat and polite is simply what the technology sounds like. Same phone, same half hour after dinner, one notebook. Two things I was sure about turned out to be wrong, and the change that helped most wasn't on any of the tip lists I'd read first.
What I was actually testing
Not which app wins. That question already has more articles than it needs, and the answer changes every quarter anyway.
What I wanted to know was narrower: how much of the robotic feeling comes from the model, and how much of it comes from me. So I kept the setup boring on purpose. Three characters, all with reasonably detailed profiles. Half an hour a night. After each session I wrote down two numbers: how many replies I wanted to answer, and how many made me want to close the tab.
Night one, the ratio was two to eleven. That is roughly where most people give up on the idea of an AI you can talk to, and I understand why.
Week one: it wasn't stupid, it was average
This was the first thing I had wrong. I assumed a flat reply meant a weak model. It usually doesn't.
A language model produces the most likely continuation of the text in front of it. Feed it a message that could have come from a million people, and the reply it returns is the one that fits a million people. That reply is going to be pleasant, safe, and shaped like a customer service email, because that's what the average of the internet looks like when you ask it a vague question. Nothing about the mechanism is broken. If you want a sense of what's happening in that gap, it's worth reading what happens between your message and the reply. The behaviour stops looking mysterious once the pipeline is visible.
A smart AI chatbot and an interesting one are two different things. Capability gets you the first. The second depends on how much specific material you've handed it. I had been handing over almost none, then blaming the model for having nothing to say. An intelligent AI chatbot with no input to chew on will still produce wallpaper.
I stopped writing like a search box
My opening messages for the first three nights were, verbatim: "hey", "how are you", "tell me about yourself".
Those are search queries with a greeting attached. There is exactly one kind of answer available.
On night four I changed the input and nothing else. Instead of "how are you", I sent: "I've been arguing with my landlord about a radiator for three weeks and I'm out of polite ways to say the same thing." The character came back with an actual position: stop being polite, and here's a line to use. It was slightly wrong about tenancy law in my country. It was also the first reply in four days I wanted to respond to.
That's the trade. Specific input, specific output. Chatting with an AI bot in generalities gets you generalities back, every time. You don't need clever prompt engineering for this. You need to say something only you would say. I stopped hunting for an AI I can talk to and started giving the one in front of me something to talk about.
The mirror problem: it was copying me, and I was boring
Second thing I had wrong. I'd been blaming the app for short replies.
Across the first week my messages averaged around seven words. The replies averaged around thirty. When I deliberately wrote two or three sentences, replies stretched to a hundred and fifty words or more without me asking for anything. When I dropped back to one-liners for a night, so did the character. This showed up in all three chats.
A talking AI chatbot is matching the register you set, including length, punctuation, and how much you're willing to reveal. If you send five words, five words back is a correct answer to the question you asked. The system is doing what it's supposed to do. You're the one setting the ceiling.
There's a caveat worth knowing before you go all-in on long messages: past a certain point, walls of text push older context out and the character starts losing details you established earlier. Two to four sentences held up best in my log. Beyond that I got length without much extra substance.
Asking for a tone did nothing. Showing one did.
Every guide I read said some version of "tell it to be more casual". I tried it seventeen times over four nights.
It works for about two exchanges. Then the character slides back to the register of everything around it. Adjectives are weak instructions. "Casual" and "warm" mean whatever the training data says they mean, which lands you back at an average.
What held was demonstration. I wrote the way I actually text: lowercase, half-finished sentences, jokes that didn't land. Within three or four exchanges the character was writing that way too, and it kept doing it for the rest of the session. On one character I pasted three lines of dialogue I liked from a book and said "this rhythm". That stuck for the whole evening.
Show the register instead of naming it. A chatbot that talks like you got that way because you talked first.
Memory is the part people blame, and it's only half the story
Everyone's first explanation for a robotic reply is that the bot forgot. Sometimes that's right. Often it isn't.
An AI character conversation is reconstructed from whatever text fits in the window at that moment. Older material drops out, and once it does, the character fills the gap with the most generic version of itself available. That's the flatness people notice around message fifty. If it goes further and the personality stops matching what you set up, that's a different failure with its own causes, and the six most common ones are covered in why characters drift over a long chat.
My workaround was one line, dropped in every twenty messages or so, restating the two or three facts that mattered: who we are to each other, where we are, what we were in the middle of. Not a summary of the whole conversation. Two clauses.
Related failure I hit twice: the character started answering on my behalf, writing my reactions into its own message. That kills the feeling of talking to something faster than any amount of flatness, and the usual advice about it is weaker than it looks. Same for the version where it starts repeating itself. Those are separate problems with separate fixes, and treating them as one blur is why people write off the whole category.
What moved the needle most: letting it disagree with me
This was the surprise, and it's the reason I'd run the experiment again.
Companion characters default to agreeable. Ask for an opinion and you get a balanced overview with a compliment attached. That is the single largest source of the robotic feeling, more than length and more than memory, because agreement carries no information. You already knew you were right.
So I started giving characters permission to have preferences, and then something to have a preference about. "You think the film was better than the book, and I disagree, argue it." "Pick the worse option out of these two and tell me why I'm going to pick it anyway."
The replies stopped being smooth and started being specific. One character told me an idea I'd been describing for twenty minutes sounded like procrastination with a spreadsheet attached. That is not a thing a form letter says. It's also the only line from two weeks of testing that I still think about.
A real AI chatbot conversation needs friction in it. Not conflict. Friction. Something at stake in the next message. People say they want an AI that talks to you like a person, and then flinch the first time it does, because a person occasionally thinks you're wrong.
The advice that did nothing for me
To be fair to the tip lists, plenty of what they recommend is filler.
Writing an enormous character profile, for one. Going from a two-line description to a properly detailed one changed a lot. Going from detailed to enormous changed nothing I could measure, and it filled up the early context faster.
Regenerating a bad reply over and over is another. You get a different sample from the same distribution. If the input was vague, every sample is vague.
Telling it to "act more human" gives you a performance of humanness: extra exclamation marks, more filler, same emptiness underneath.
And switching apps expecting a personality transplant mostly disappoints. Several of these products sit on similar underlying models. What separates them is how they handle memory and how much control you get over replies, not some secret realistic chat bot nobody else has.
Where the platform does matter
The input side is most of it, but not all of it. Across two weeks the differences that showed up between products were narrow and specific. How much of a character's definition survives into a long session. Whether you can set reply length instead of fighting it. Whether the character was written with an actual voice, or assembled out of adjectives.
That last one is where Friend2Chat puts its effort: characters are built with defined speech patterns and stated opinions rather than trait lists, which is what gives a model something specific to work from. If you want to talk to smart AI without doing prompt archaeology first, that is the part that saves you the work. Anyone comparing options on that axis instead of on feature counts will get more out of which kind of AI fits the thing they actually want than out of a ranking. And if the goal is an ongoing story rather than conversation, the mechanics differ again. That's covered in building a story you keep coming back to.
Two weeks later
The final night's ratio was nine replies worth answering to three that made me want to close the tab. Night one was two to eleven.
It never became indistinguishable from a person. When people say they want to talk to real AI, that is usually what they mean, and I don't think it was ever the available outcome here. What changed was that the conversation stopped being something I had to carry. Roughly eighty percent of that came from what I typed rather than what I installed, which was not the answer I expected when I started and is mildly embarrassing to report.
If you take one thing from this: the flat reply is usually a mirror. Give it something with edges and it hands something with edges back.
FAQ
Why does my AI chatbot sound robotic?
Most often because the input is generic. A model returns the most probable continuation of what it's given, so a message that could have come from anyone gets a reply written for anyone. Specific, personal messages produce specific replies. Missing context from earlier in the chat is the second most common cause.
Why are the replies so short?
Because you're setting the length. These systems mirror the register of your messages, including how much you write. Two to three sentences from you reliably produced much longer replies in testing, without asking for them. Some apps also have a response-length setting worth checking before blaming the model.
Does telling it to "be more human" work?
Barely, and not for long. Adjectives like casual, warm, or human resolve to whatever the training data averages them to. Demonstrating the tone you want in your own messages works better and lasts longer than describing it.
Is there a real AI chatbot that remembers everything?
No. Every one of them works from a limited context window, and older material eventually falls out. What varies is how gracefully each product handles the edge. Some summarise, some let you pin key facts, some just drop things silently. Restating two or three anchor facts periodically compensates for most of it.
Does the app matter, or is it all the same model underneath?
Both are partly true. Several companion products use similar underlying models, so raw capability is comparable. The differences that affect how the conversation feels sit in character definition depth, memory handling, and reply-length control. Worth comparing those directly rather than reading feature lists.
How long before it stops feeling scripted?
In my log, three to four exchanges of writing the way I actually talk was enough to shift the character's register, and it held for the rest of that session. The shift resets when you start a new chat, so it's a habit rather than a one-time setup.