Week One: 900 Questions, Murderer Tells, and the Birth of Falco

Week One: 900 Questions, Murderer Tells, and the Birth of Falco

One week in, and I'm still processing how much we've learned.

The Numbers

10 players. 26 games. 900 questions. Average session: 1 hour.

Not massive numbers, but enough to surface patterns I never expected. Some players went full roleplay - adopting detective personas, building elaborate theories, treating suspects like real people. Others took the analytical route: clinical questioning, fact-gathering, motive-mapping.

Some of the characters generated this week

What fascinated me most? The split between motive hunters and method hunters. Some players focused entirely on why - who had reason to kill? Others obsessed over how - what evidence proves capability? Both approaches work, but they create completely different gameplay experiences.

The Murderer Tell

Mid-week, a player messaged: "I can spot the murderer now."

That's... not supposed to happen.

Turns out our playtesters had noticed subtle differences in the AI agent prompts for murderers versus innocent suspects. Once you knew what to look for, you could identify the killer in other games without asking a single question. Classic tell.

We fixed it immediately—normalised the prompts, removed the signal. No more easy shortcuts.

The Hallucination Problem

We hit a bigger issue: AI hallucination between the original mystery setup and the suspect agents.

The generator would create a scenario - let's say, a poisoned wine glass at a dinner party - but when suspects answered questions, they'd reference details that never existed in the original prompt. The wine glass became a champagne flute. The dinner became a cocktail hour.

Small discrepancies, but they broke immersion.

The fix: enrichment passes and regeneration of the summary. We now run additional processing to ensure the mystery setup and all suspect backstories stay consistent. The side effect? Much richer, more descriptive scenarios. The intros now actually lead into the suspects and initial evidence instead of feeling disconnected.

The Difficulty Crisis

By Thursday, I noticed something troubling.

Our average solve time had dropped from 30-40 questions to just 16-20 questions. Players were cracking cases too easily. The enrichment passes that fixed hallucinations had inadvertently made the mysteries more transparent.

I realised we needed difficulty tiers, not just one-size-fits-all mysteries.

So we built Falco - our second game generator (named after a fictional detective, like all our generators). Falco uses different generation techniques to create harder cases with more misdirection, red herrings, and complex motive webs. Now we can offer both easier and harder games depending on what players want.

Architecture overlay of the generators

We'll do a technical deep-dive on the generator architecture soon. It's one of the more interesting engineering problems we've solved.

What Worked

Infrastructure: flawless Running on Heroku with both Anthropic and ChatGPT as model backends. Everything scales smoothly. Multi-model architecture paying off - we can optimise each function independently.

Player engagement: **strong**. One-hour average sessions mean people are genuinely invested in solving cases, not just clicking through.

What We Shipped This Week

Based on your feedback, we've already rolled out:

- Improved onboarding guide to help new detectives get started - Better evidence popups that make clues more discoverable - Social sign-ins for faster access

What's Next

Immediate priorities:

- Mobile optimisation (you keep asking, we're listening) - Fine-tuning Falco to generate more challenging mysteries with better misdirection

Longer-term: We're watching play patterns closely. The motive vs. method split suggests we might want to design mysteries that reward different investigation styles.

Thank You

This week wouldn't have been possible without our playtesters. You found the murderer tell. You reported the hallucinations. You told us what worked and what didn't. You're not just playing Improbable Remains—you're building it with us.

Thank you for being part of this from the beginning. Your feedback is shaping what this becomes.

As a reminder, during the launch beta period, all games are free to play, and beta testers will also be rewarded with credits at the paid-launch.

James
James

Engineer by day, indie game developer by night. Built Improbable Remains to combine AI technology with classic murder mystery gameplay. Python coder with a passion for creating intelligent, interactive experiences that challenge players to think like detectives.