My story
I’m a Conversational AI Systems Specialist focusing on long-horizon human–AI interaction, conversational AI evaluation and human-centred AI system design. My work explores what happens when people and conversational AI systems remain in conversation for weeks or months rather than minutes — long enough for memory, adaptation and entirely new interaction patterns to emerge.
While much of today’s conversational AI evaluation focuses on isolated prompts and short benchmark conversations, my research examines conversational AI in naturalistic, long-running settings. I’m interested in how conversational AI systems accumulate context, develop shared shorthand, retain corrections, recover from mistakes and evolve over time. These behaviours become increasingly important as AI moves from answering individual questions to becoming an ongoing collaborator.
I came to that question the long way round.
I studied Computer Science in Poland, specialising in systems design. My career began in the games industry, testing software across the USA, Germany and Spain before moving into localisation for Sony PlayStation, Electronic Arts and Jagex on franchises including Uncharted, God of War, Resistance, Ratchet & Clank and The Sims 2. Later I managed Quality Assurance at Cambridge University Press, founded and ran my own business, and today lead automation-first, AI-augmented localisation at LiveScore.
The industries changed. The systems perspective never did.
Across games, publishing and AI, I’ve always been interested in the same question: how do complex human-technology systems behave, where do they fail, and how can they be designed to work better?
In 2023 I began asking a question that very few people were studying directly: what happens when a human and a conversational AI remain in the same conversation for hundreds or even thousands of turns?
As conversations grew longer, I observed recurring patterns that conventional AI evaluation rarely captures. Conversations developed shared shorthand. Corrections persisted across time. Earlier interactions continued to influence later reasoning. The conversation itself became an evolving system whose behaviour could not be explained by any single prompt.
Rather than treating those observations as anecdotes, I began documenting them systematically through longitudinal studies of naturalistic human–LLM interaction. That work led to my first research framework and my first academic paper submission: the provocation Nobody Knows What Really Happens at Turn 1000: Who Is Really Speaking in Human–LLM Interaction? — which scored 8/10 from two knowledgeable reviewers and laid the foundation for my subsequent research. It was followed by the publication of the High-Coherence Interaction State (HCIS) framework, published with a DOI.
To support that research, I built my own persistent conversational AI research platform — not as a product, but as an experimental environment for studying conversational AI over extended periods. It combines persistent memory, longitudinal interaction, multi-agent architectures and extended reasoning, allowing me to observe how conversational AI systems change across months of real-world use.
Alongside that platform I developed original evaluation methods for conversational AI, designed to measure behaviours that traditional benchmarks often overlook. These include engagement stability, conversational drift, context weaving, correction persistence, interaction coherence and other longitudinal characteristics of human–AI collaboration.
Over the past several years I have spent thousands of hours conducting, analysing and designing long-horizon conversations with frontier language models. That sustained observational work forms the foundation of everything here, connecting theoretical frameworks with practical AI system design and real-world conversational behaviour.
Alongside the practical work, I’ve continued to build formal expertise through Human–Computer Interaction for AI Systems Design at the University of Cambridge, AI Governance with the IAPP, and professional study in AI evaluation and online child safety — the full list is here.
I’m also the author of The Ultra-Long Horizon, a forthcoming book exploring what really happens inside very long conversations between humans and conversational AI — and why many of the most important behaviours only emerge after hundreds or thousands of conversational turns.
My first job was making sure a game spoke naturally in thirty languages. Today I’m asking a different version of the same systems question:
What is a conversational AI system really saying after one thousand turns — and how much of that behaviour belongs to the model, the human, or the interaction they’ve built together?
Sounds interesting? Contact me.