AI Chat That Sends Pictures: How In-Chat Photos Actually Work
"She sent me a photo" is the feature every AI chat app now advertises, and in most of them the photo is of someone else. Right hair, roughly the right outfit, a face that changes a little every time. Ask for a second picture and the drift is obvious.
This is an explainer on what actually happens when an AI chat sends pictures: the three ways apps do it, why two of them produce lookalikes, what it takes to draw the actual character from what she just said, and what that costs. We build the third kind, so the last section is biased and says so.
The Three Ways an AI Chat "Sends Pictures"
1. A gallery
The character has a fixed set of images uploaded by whoever made her. The chat picks one that roughly matches the mood. It is consistent because nothing is generated, and it is useless the moment the story goes somewhere the gallery did not anticipate. Character platforms with user-made bots mostly work this way, with a generated option bolted on later.
2. A general image model handed a description
The app takes the reply, or your request, turns it into a prompt ("red-haired woman, black suit, office, smiling") and sends it to a general image generator. This is what most companion apps and uncensored chat platforms do, and it is why the pictures look like a lookalike: the model has never seen her. It has seen millions of red-haired women. Every picture is a new guess at the same description, so the face drifts, the outfit changes, and details the reply mentioned get dropped.
Reviews of Talkie, PolyBuzz and CrushOn in 2026 all describe the same symptom - "face drift" between images - and Character.AI's in-chat image feature, launched in March 2026, is the same class of generic generation with a nicer interface. Perchance's built-in generator is the honest budget version of it: small, free, and not pretending to know who she is.
3. A model trained on the character
The image model was trained on her specifically, so it knows what she looks like from every angle, in her canonical outfit and out of it. Given the same description as in option 2, it draws her. This is what anime image generators with named character catalogs do, and it is what makes "she sent a photo" mean what it says. The cost is real: someone has to train and maintain one model per character, which is why none of the general chat platforms do it.
"From Her Reply" Is the Part That Matters
Even with a trained model, there is a second question: what is the picture of? Two designs exist.
You describe the picture. You open a separate prompt box and type what you want to see. Works, but it pulls you out of the story every time, and the picture has no connection to what she just said.
The picture is grounded in her reply. When she takes a photo in the story - she says she is doing it, and her reply says what the photo shows - the app draws exactly that, from her words. If she wrote that she is on the bed with the lamp on and her jacket on the chair, that is the photo. Nothing she did not say gets invented; nothing she said gets dropped. The same works for any reply that describes something visible: you press one button and the scene she just narrated is drawn.
The second design has two consequences worth knowing. First, the writing has to carry the picture, so a good roleplay model is told to name concrete visible details during actions - which also happens to make the prose better. Second, you are the camera. She is drawn; you are not painted in, and the picture is from where you are standing. It is a photo she sent you, not a still from a film about the two of you.
Continuity: Why the Second Photo Is the Hard One
The first photo is easy. The second one is where lookalike systems fall apart, because each generation starts from zero: outfit resets to default, the room changes, the pose forgets what happened.
A chat that sends pictures properly has to carry state between them. Outfit, state of undress, position, place - locked when established, carried into the next picture, and updated when the story changes them. If her jacket came off two scenes ago, it stays off until the story puts it back. This is a hard problem and mostly invisible when it works, which is exactly why it is worth checking: ask for two photos five messages apart and see whether the second one remembers the first.
What It Costs, Honestly
Generating a picture takes GPU time, and there is no way around paying for it - the only question is whether you pay per picture or through a subscription that hides the count. A few reference points as of September 2026: Candy AI bundles a monthly token allowance into a subscription and charges tokens per image; CrushOn puts images behind paid tiers with monthly caps; PolyBuzz blurs images on its free tier and unblurs by subscription level.
Per-picture pricing is the honest version because you can see what a photo costs and stop whenever you like. The thing to check is what happens when a picture fails to generate: a refund, or a lost credit.
Who Does What in 2026
| Service | How pictures are made | Is it actually her? | Where it lives | Continuity between pictures |
|---|---|---|---|---|
| Character.AI (Imagine) | General generation, partly paid | No | Mobile app | No |
| Talkie, PolyBuzz | General generation | No, face drifts | App | Limited memory |
| CrushOn, SpicyChat | General generation, paid tiers | No | Web | No |
| Candy AI, companion apps | Persona-consistent general model | Consistent OC, not a real character | App/web, subscription | Some |
| Perchance | Small free generator | No | Browser | No |
| goongen.ai Roleplay (in build) | Model trained on each character, drawn from her reply | Yes | Web, encrypted | Outfit, place, position carried |
None of the services above train a model per character, and none of them encrypt what they store; both of those are the expensive choices, which is why they are rare.
Where We Fit
We are the team behind goongen.ai, so read this as the biased part.
Roleplay is in build and not open yet. The AI roleplay page shows what it will look like. What is live today is the image side those photos come from: 1,400+ anime characters in our catalog, each drawn by her own trained model, no filters beyond the two legal hard lines.
The chat is being built as design 3 with pictures grounded in her reply:
- When she takes a photo in the story, it appears in the chat, drawn from what her reply literally says by the model trained on her. Not a lookalike.
- Generate scene on any reply that describes something visible, one button, no separate prompt box.
- Outfit, state of undress, position and place carry from one photo to the next.
- You are the camera. She is drawn; you are never painted in.
- 5 credits per delivered photo, refunded if it fails to render. Replies cost 1 credit each. Packs from $0.99, crypto only, no subscription.
Every photo is locked before it is saved so only your password opens it; we never see the password, because the login checks it without it leaving your device. We could not look at your pictures if we wanted to. There are no ads and nothing is mined - credits pay for the GPU time a picture takes, and that is the whole business model.
The honest tradeoffs: not shipped yet; anime characters first, realistic ones later; crypto only; and if you lose your password and the backup file, the pictures are gone with everything else.
Closing
An AI chat that sends pictures is only as good as two things: whether the model knows who she is, and whether the picture is of what she just said. Galleries and general generators fail the first; separate prompt boxes fail the second. Ask for two photos five messages apart and you will know which kind you have.
The models that draw our characters are live in the anime generator today - the demo gives every account 5 free images, no card - and the AI roleplay page is where you will hear when she starts sending them herself.