ARISTOTLE BETA
An AI co-scientist built for researchers trained to question everything.
Role · Product Designer 0 → 1
Timeline · Oct 2025 - Feb 2026
Team · 2 designers + 3 engineers
THE PROBLEM
Scientists don't trust AI. And they shouldn't have to.
Most AI tools answer every question with the same flat confidence, no matter how solid the evidence actually is.
No way to see how confident an answer should be
Reasoning happens in a black box, nothing to interrogate
Built for casual use, not rigorous methodology
THE VISION
To close that gap, I studied how researchers actually evaluate evidence and rebuilt the interface around it.
Transparency is the product.

WHERE WE STARTED
Just ask and Aristotle chooses the best model for you.
The first version routed every prompt automatically across three models, Explore, Generate, and Verify. No sidebar. Clean canvas. Minimal footprint. The logic was simplicity: reduce friction, hide complexity, let researchers just ask.
It worked until it didn't.
WHERE IT BROKE
A black box, for people trained to question everything.
Researchers couldn't see why Aristotle responded the way it did. When a response felt off, there was nothing to interrogate. At the same time, the clean canvas hit its limit, Double Check needed space, source previews needed space, reasoning traces needed space. Surfacing it all inline turned the canvas into noise.
Two things had to change.
THE PIVOT
Using model settings to communicate intent.
We gave researchers control of model selection, not as an afterthought, but as a way to choose a research posture. We introduced the sidebar as a dedicated reasoning layer, separate from the canvas, visible without competing with it. It became the foundation for everything that came after.
From black box to trusted.
No proper design system. No precedent. Three months to scale.
I rebuilt model selection and the reasoning layer around visibility instead of automation, and shipped to qualified researchers in February 2026.
IMPACT

1,400+
researchers

1,000+
institutions
Under 3 months.
WHAT I SHIPPED
X1 model responses · Projects · Double Check · System Thinking · In-line definitions · Source previews
Live to qualified researchers — February 2026
Four models, one mental model.
X1 Spark for the messy early question. X1 Search for literature depth, with a visible reasoning trail. X1 Verify for pressure-testing its own answer. X1 Instant for when a lighter touch is enough.
X1 SPARK
Hypothesis generation through self-skeptical reasoning.
Scientists don't always arrive with a fully formed question. Spark is built for that earlier, messier phase. It surfaces cross-disciplinary directions that wouldn't come up in a standard literature search, and it does it by challenging its own outputs as it goes.
X1 SEARCH
Literature synthesis at depth. Search structures evidence across sprawling research areas and makes every claim traceable.
The sidebar carries the reasoning trail: steps like "Scoping the Question," "Structuring the Approach," "Checking the Numbers," so researchers can follow the methodology, not just read the conclusion.
X1 VERIFY
Skepticism as the primary mode. Verify challenges its own reasoning before delivering an answer.
The sidebar surfaces this visibly: Answer Confidence, Source Verification, Mathematical Validation, Causal Reasoning — each as a discrete, inspectable layer. For scientists who need to know not just what Aristotle concluded, but whether it held up.

X1 INSTANT
Focused questions, fast. Source-backed answers in seconds, without the overhead of a full research workflow.
The interface stays lean: no expanded sidebar, no reasoning trail, because the use case doesn't need it. Instant trusts the researcher to know when a lighter touch is enough.

DOUBLE CHECK
AI that sounds confident is a liability in science.
Double Check automatically verifies factual claims inline and flags weak assertions before a researcher builds on them.
Three states: Verified, Uncertain, Flagged, each with supporting evidence visible on expand. The design problem was making this feel like a tool for rigorous researchers, not a disclaimer for a nervous product.
INLINE DEFINITIONS
Dense scientific language creates friction across disciplines.
Inline definitions surface context exactly where it's needed, without breaking reading flow.
The key decision was keeping them genuinely inline: no modal, no navigation away. Present when needed, invisible when not.

PROJECTS
Research isn't a single query.
Projects create dedicated memory spaces where related work stays together: chats, files, context, so Aristotle maintains coherent recall across sessions.
The structure mirrors how researchers actually work: sustained lines of inquiry, not one-off searches.

WHY THIS PAPER
Every citation carries an implicit question: why does this source matter here? Why This Paper answers it directly, surfacing the goal, the contribution, the methodology, and the limitations for each source.
It closes the gap between AI output and primary source verification without making the researcher leave the interface to do it themselves.
Beta was step one. Trust made it possible.
Autopoiesis is building AI that can do science autonomously, not an assistant that waits to be asked, but a system with the judgment to sustain inquiry under uncertainty on its own.
Every decision in Beta, the sidebar, the model picker, the transparency features, was scaffolding for that system. Beta proved skeptical scientists could trust an AI co-scientist. Everything after this builds on that.
Live to qualified researchers — February 2026





