Menu
Projects

ProductDesktop AI research platform2023–2026

Ubik Studio

A desktop-native, local-first AI research platform. Three and a half years building the human side of agentic research, before it had a name.

The workspace in its final build: the file explorer and an indexed-four-minutes-ago context pill on the left, a source paper in the middle with every claim the agent drew highlighted in place, evidence cards queued for review along the bottom, and the agent working through a four-task plan on the right

The workspace in its final build: the file explorer and an indexed-four-minutes-ago context pill on the left, a source paper in the middle with every claim the agent drew highlighted in place, evidence cards queued for review along the bottom, and the agent working through a four-task plan on the right

2023–2026 · co-founded · the public site and builds are retired; the test log and the subreddit are what remain in the open.

In simple terms

Ubik Studio was a desktop program that worked like a research assistant: the AI did the reading and the drafting, and a person stayed in charge of every judgment call.

You pointed it at an ordinary folder of PDFs and papers on your own computer. It read and indexed them locally, searched across dozens at once, and pulled out what mattered. The rule it enforced was simple: the AI could not state a fact or write a sentence unless it could attach a direct quote and an exact page number from your own files.

When something needed a human decision, the agent stopped mid-task and said so. You got a review queue and approved, rejected or edited each claim before it reached the draft.

The impact. Most AI writing tools of the era tried to replace the researcher and shipped a wall of text nobody could check. Ubik did the opposite: it took the tedious half of the work and made the person verifying it the point of the product, which is the pattern the industry has spent the years since rebuilding.

From the archive. Built to Learn, Not to Help · Amplifying Human Intelligence · The Ubik Portals Bible

Why this was ahead of its time in 2023

In 2023 the state of the art in public was a single chatbot in a text box. It hallucinated freely, could not cite anything reliably, and left you copying snippets back and forth by hand. Ubik was already doing five things the field would spend the next several years arriving at.

  • Multiple specialised agents, not one general chatbot

    Then you talked to one model that tried to do every part of the job at once.

    Ubik split the work across sub-agents, a PDF reader, a literature researcher and a drafting agent, and let you assign a different model and reasoning depth to each. A cheap fast model ingested text; an expensive one wrote.

  • Citations enforced in code, not requested in a prompt

    Then early chat-with-your-PDF tools were notorious for inventing quotes and giving page references that did not exist.

    Ubik made it a programmatic rule. No exact quote and page number meant the claim could not be written at all, and every drafted line stayed linked to the highlighted source.

  • Human-in-the-loop as architecture rather than a confirmation dialog

    Then you got either full automation or a tedious one-prompt-at-a-time conversation.

    Ubik had a dedicated review queue and a Human Needed status. The agent paused at judgment calls, queued the evidence, and waited to be approved, rejected or corrected in batches.

  • Local-first, on your own machine

    Then nearly everything required uploading confidential research to somebody else’s cloud.

    Ubik ran on the desktop against your real file system. Pointing it at an existing folder indexed it in place, with no proprietary silo to move your data into.

  • Reading many long documents at once

    Then context windows were often four to eight thousand tokens, which made cross-referencing several papers essentially impossible.

    Ubik let you mention a dozen papers in one prompt and read them in parallel against a working draft instead of summarising them one at a time.

While most companies that year were putting a wrapper around a chatbot, this was already the interaction design, the verification safeguards and the multi-agent workflow the rest of the field would go on to build.

What Ubik was

Ubik Studio was a desktop research environment where AI agents did the gathering, reading, and drafting — and the human stayed in the loop at every point of judgment. A researcher opened a folder, and it became a workspace: sources indexed into a local context engine, agents searching the literature and reading PDFs in parallel, documents drafted with citations that traced back to real pages — and a review queue standing between every consequential agent action and the workspace it wanted to touch.

The thesis was written directly into the agents themselves. From the writing agent's system prompt:

“Your job is not to replace human thinking — it is to amplify it. Optimize for the loop: you draft, the human refines, you incorporate, the human approves. Intelligence is maximized not when either side works alone, but when the handoff between AI and human is so seamless it feels like one mind thinking.”

That sentence was written years before “human-in-the-loop” became an industry talking point. Ubik spent three and a half years trying to actually earn it.

The product

A desktop app where any folder became a research workspace. Sources went in — papers, PDFs, webpages captured by a companion browser extension — and were indexed locally. Agents searched the literature, read sources in parallel, and drafted documents where every claim traced back to a real quote on a real page. Everything an agent wanted to do of consequence passed through a human first: sources were approved before they entered the workspace, drafts paused for judgment where judgment was needed, and nothing was written that couldn't be cited.

The human-in-the-loop architecture

The part of Ubik I'm proudest of is that human control wasn't a confirmation dialog bolted on at the end — it was load-bearing architecture. Agent actions were approved in batches, not rubber-stamped one toast at a time. Every review decision was recorded in an auditable trail you could revisit after the fact. Agents could stop mid-document and ask for human judgment exactly where it belonged, at a depth you could dial from rough scaffold to polished draft. And the rule I still think about most: if there was no evidence to cite, the agent didn't get to write the claim.

The product, in motion

Seven silent recordings of the last build, March 2026, each one looping on its own.

A folder becomes a workspace

1:25

Files stay on machine so Ubik Agents work with context locally.

A folder becomes a workspace — No import step and no database: point Ubik at a directory of PDFs and it indexes them in place, into a local context engine that reports how stale it is in the status bar. The explorer is the file system, so the workspace is still just a folder when you close the app.

No import step and no database: point Ubik at a directory of PDFs and it indexes them in place, into a local context engine that reports how stale it is in the status bar. The explorer is the file system, so the workspace is still just a folder when you close the app.

Twelve papers into one prompt

1:15

Delegate busy work so you can focus on what matters.

Twelve papers into one prompt — Sources are @-mentioned rather than uploaded. A dozen papers named in a single sentence, and the agent reads them in parallel against the draft rather than summarising them one at a time. The model picker sits beside the send button because which model reads your sources is a decision worth leaving with the researcher.

Sources are @-mentioned rather than uploaded. A dozen papers named in a single sentence, and the agent reads them in parallel against the draft rather than summarising them one at a time. The model picker sits beside the send button because which model reads your sources is a decision worth leaving with the researcher.

Search that scores its own results

0:53

Search academic databases and pull papers into your workspace with a click.

Search that scores its own results — Agentic literature search across 146 results and 8 searches, then the Result Explorer: every candidate scored on peer review, methodology, recency and relevance, with an agent analysis under each and a preview before anything is accepted into the workspace.

Agentic literature search across 146 results and 8 searches, then the Result Explorer: every candidate scored on peer review, methodology, recency and relevance, with an agent analysis under each and a preview before anything is accepted into the workspace.

A note is its evidence

1:16

Annotate files, build bibliographies, and work with cited output.

A note is its evidence — Notes are typed — highlight, key point, claim, evidence, definition, methodology, limitation — and opening one shows the supporting quotes with page numbers behind it. There is no note in Ubik that is only an assertion; if the quotes are not there, the note does not get written.

Notes are typed — highlight, key point, claim, evidence, definition, methodology, limitation — and opening one shows the supporting quotes with page numbers behind it. There is no note in Ubik that is only an assertion; if the quotes are not there, the note does not get written.

Human Needed, and the review queue

0:30

Stay confident with human-in-the-loop features that amplify intelligence.

Human Needed, and the review queue — The part I am proudest of. The agent stops mid-document, states the judgment it needs, and waits — skip or submit, counted in the status bar as blocking. Beside it the review queue holds every claim the run produced, each one accepted or rejected on its own, with a jump straight to where it would land.

The part I am proudest of. The agent stops mid-document, states the judgment it needs, and waits — skip or submit, counted in the status bar as blocking. Beside it the review queue holds every claim the run produced, each one accepted or rejected on its own, with a jump straight to where it would land.

Model control, per subagent

0:50

Route to frontier models, experimental models, and local models. You decide what runs where.

Model control, per subagent — The PDF reader, the writer and the researcher are separate agents, and each one takes its own model and its own reasoning effort. Toolsets switch off individually. A researcher who wants a cheap model reading PDFs and an expensive one writing the draft can have exactly that.

The PDF reader, the writer and the researcher are separate agents, and each one takes its own model and its own reasoning effort. Toolsets switch off individually. A researcher who wants a cheap model reading PDFs and an expensive one writing the draft can have exactly that.

Hopper, the capture extension

0:29

A browser companion. Browse any paper or webpage and hop it straight into your workspace with full metadata.

Hopper, the capture extension — The companion browser extension. Hop a page and it lands in the workspace sources folder, verified and attributed, syncing to the desktop app when it is connected. The web half of a local-first tool: the browser is where sources are found, and the file system is where they live.

The companion browser extension. Hop a page and it lands in the workspace sources folder, verified and attributed, syncing to the desktop app when it is connected. The web half of a local-first tool: the browser is where sources are found, and the file system is where they live.

Three longer runs from the 2025 build2:33 – 3:26

Whole sessions rather than single capabilities. The interface is a year older and the loop is the same one.

Workspace walkthrough — The full loop, end to end: ask the agent what’s working in a draft, and it reads the essay against the source PDFs, writes analytical notes, and queues every one for review — accepted or rejected claim by claim, with the evidence panel surfacing the supporting quotes as you go.
Workspace walkthrough3:26

The full loop, end to end: ask the agent what’s working in a draft, and it reads the essay against the source PDFs, writes analytical notes, and queues every one for review — accepted or rejected claim by claim, with the evidence panel surfacing the supporting quotes as you go.

Cross-source synthesis — Five papers @-mentioned into a single prompt — “find the commonalities” — while the Context Engine indexes the workspace live. The agent reads the peer-reviewed PDFs side by side and builds a synthesis grounded in all five sources, not a summary of one.
Cross-source synthesis2:33

Five papers @-mentioned into a single prompt — “find the commonalities” — while the Context Engine indexes the workspace live. The agent reads the peer-reviewed PDFs side by side and builds a synthesis grounded in all five sources, not a summary of one.

The workspace library — The file explorer as a research surface: PDFs, documents, saved searches, and folders in one indexed library. Then a multi-note agent run — grep across source bundles, evidence distributed in bulk, three notes finalized — with every artifact landing in the explorer as it’s produced.
The workspace library2:36

The file explorer as a research surface: PDFs, documents, saved searches, and folders in one indexed library. Then a multi-note agent run — grep across source bundles, evidence distributed in bulk, three notes finalized — with every artifact landing in the explorer as it’s produced.

The marketing site

The product was not the only surface I owned. Ubik's marketing site was mine end to end: the argument it opens with, the copy, the structure, the paintings commissioned for each capability, and the build. It was written in the order the product works. The problem first, then how Ubik answered it, then twelve fields of use cases with worked examples in each, and a models page that said what ran on the free plan.

The site is retired with the rest. The Wayback Machine kept a copy from April 2026, and these are its pages; the archived site still opens.

The engineering

A large multi-package system: a desktop app, a web gateway, cloud agent deployments, and a browser extension, with a local-first storage model underneath it all. 1,038 commits from September 2023 to May 2026 — with the design and research that preceded the first commit, about three and a half years of my life.

Electron · Next.js · TypeScript · Python · local-first

My role

I co-founded Ubik and led product design end to end: the workspace model, the review surfaces, the evidence and citation UX, the Human Needed grammar, and the copy and interaction conventions across every surface. I built front-end throughout, and ran the user research cycles — interviews, behavioral observation, session replays, and the team test log that documented them in public.

On the agent side I owned the system prompts, skills, and custom datasets — and I designed and built the custom evaluation framework and ARC eval suite we used to train, tune, and regression-test our agents and the agent-orchestration systems that coordinated them. That evaluation work is what actually moved the product: measurable gains in output accuracy, answer quality, and real-world usability — not benchmark numbers in isolation, but whether a researcher could trust and use what came back. It's the part of Ubik least visible in a screenshot and the part that mattered most to the results.

The design board

Ubik never had a design team, a seat of anything for everybody, or the time to keep a spec in sync with itself. What it had was Excalidraw files that nobody ever closed. This is one of them, running from 2024 into 2025, left exactly as it was.

It is a snapshot rather than the archive. Plenty of boards came before it, and the 2023 and 2024 ones from the Fiig years and the ed-tech work sprawl further than this. I picked this one because it is a fair picture of how I actually think while a product is still being decided.

It worked because it refused to be one thing. The same board carried landing page explorations, user story wireframes, screenshots of the running app with corrections drawn straight over them, and plain notes to whoever opened it next. A sketch on the left, the decision written beside it, and underneath that the file path it applied to. That middle layer, looser than a spec and more durable than a conversation, is where most of this product actually got decided.

It is messy in places and I have left it messy. Areas trail off, copy is labelled not solid, and one region is simply the words ICONS NEEDED above a list of what still needed drawing. That is what a working document looks like while it is still working. Excalidraw being free is not a small detail either: in a team with scrappy limits it meant everyone could open it, and no part of how we thought sat behind a licence we were deciding whether to renew.

The Ubik Drive product design canvas: landing page explorations, user story wireframes, and review notes drawn across one board.
loading the board

drag to move · pinch to zoom

Pick an area to fly to it, or drag straight into the board.

What the board eventually taught me was when to stop drawing. Once my engineer teammate had a framework standing, it was faster to develop the flow directly in code: build the primitives, get the real UI working inside the constraints that already existed, and skip a wireframe that could only ever approximate them. That shift made me pick up a lot of new skills on the way and it changed how I design. I still open a board when a problem is genuinely unresolved, but much of what used to become a sketch now goes straight into components.

From the archive

Ubik wrote about itself while it was alive: a newsletter, a blog, internal design documents. The pieces that survive are rebuilt here in full, in their own words, as artifacts of what the company believed while it was believing it.

A closed chapter

The public site and the builds are retired, and I'm at peace with that — Ubik was a complete thing, and it stands on its own. Three and a half years of asking one question in earnest: what does it take to make an AI research tool a person can actually trust? Everything I learned answering it — about evaluation, about evidence, about where a human has to stay in the loop — I still carry. But this page is the record of the work itself, not a stepping stone to something else.