Menu
Writing

Archive · Ubik

Built to Learn, Not to Help

Archived post from the Ubik blog on kraa.

AI tools have surpassed experts, professionals, teams, institutions, really anyone in fields that require accurate citation and evidence attribution in their high-level knowledge work. At least, that is what the benchmarks say.

Legal teams, data and research teams, consultants and analysts, scholars and academics: what do individuals and teams in these fields have in common? They all work with files. Many high-level knowledge workers either have large file collections built over long periods of time or the research skills to find, analyze, and cite trusted sources found through academic databases. These kinds of field-specific, expert-level skills are unreplicatable by chatbots and platforms like ChatGPT, Perplexity, or Liner that vary in accuracy, tool access, and context awareness. All are prone to hallucination in generated text when answering one-shot prompts without context, and perform even worse when working with uploaded text files. These platforms all rank highly on published benchmarks that evaluate LLMs and AI agents. The benchmarks are the first problem:

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. [They] are not keeping pace in difficulty: LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities.

Phan et al., 2025

This quote from Humanity's Last Exam (HLE) outlines why there is a clear need for more complex, human-like benchmarks used to evaluate LLMs in their accuracy and intelligence. Humanity's Last Exam is "a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage." This "exam" is publicly available for companies and researchers to evaluate their models with, to get a better understanding of how LLMs struggle with complex tasks. While there are many different benchmarks, few focus on multi-hop questions where reasoning and problem solving guide the model rather than searching and relaying information correctly. HotpotQA is an early dataset for "diverse, explainable, multi-hop question answering," but this dataset was published in 2018, well before the powerful AI models and agents we have today. One of the dataset's own authors, Peng Qi, has since said it outright: "Stop using HotpotQA for agent research," because:

Like many question answering datasets of its day, HotpotQA is an extractive question answering dataset following the pioneering work in SQuAD, which means that the answer came directly from a substring in the Wikipedia context that supports the answer. While this leads to relatively easy-to-implement and objective evaluation metric, it leads to a format that is no longer natural for today's generative AI systems (or to real users).

Peng Qi

Nearly seven years later, Humanity's Last Exam is built to show how far LLMs have to go before being capable at complex human tasks that require intellectual agility and multi-hop reasoning. This dataset tests "structured academic problems rather than open-ended research or creative problem-solving skills," which means we aren't testing for Artificial General Intelligence (AGI), or the ability for these models to complete deep research tasks autonomously. Instead, HLE helps researchers measure technical knowledge and reasoning.

How are frontier models scoring on Humanity's Last Exam?

Chart of AI progress on Humanity's Last Exam from late 2024 through early 2026, with frontier models climbing from under 10% to around 40% accuracy
AI progress on Humanity's Last Exam, as the post charted it.

In 2025, no LLMs surpassed 50% accuracy on Humanity's Last Exam. With a title so bold this may sound like a relief at first, but these kinds of complex, multi-hop questions found in high-level research and academia aren't solvable with just a simple search and reply. HLE confirms that LLMs excel in simple, disposable tasks that require low amounts of reasoning and problem solving. For example, there are publicly available benchmarks for academic deep reasoning like GPQA and MMLU: models like Gemini 3 Pro and GPT-5 score over 90%, but the same models score below 40% accuracy when tested on the HLE benchmark.

Bar chart comparing model scores on HLE, GPQA, MATH and MMLU: the older benchmarks sit near the top of the scale while HLE bars stay near the bottom
The same models against HLE, GPQA, MATH and MMLU. The old exams are saturated; the new one is not.

When doing deep research or high-level knowledge work, humans need accurate, evidence-rich information that can be trusted and used in real, publishable work. These requirements prevent many experts and professional teams from integrating AI tools into their workflows.

SimpleQA factuality benchmark bar chart with Liner Deep Research first at 95.3, ahead of Perplexity, Gemini 2.0 Flash, GPT-4.5, Grok 3 Beta and DeepSeek R1
The SimpleQA factuality benchmark, with Liner's deep research on top.

Liner is an amazing way to quickly locate relevant papers for high-level knowledge work. It focuses on one-shot tasks that "help users discover reliable sources, generate citations, and accelerate their learning." Liner ranks first on the SimpleQA benchmark, an "OpenAI-accurate fact-checking search results" evaluation, and succeeds in delivering accurate answers with generated text that includes variations of direct quotes, citations, and links to relevant search results.

Still, Liner fails instantly when working with files added to the context window.

Let's run this simple, one-shot prompt through Ubik Studio and Liner AI. Q: Can you highlight important quotes in the file @entropy-25-00706

Liner response:

Liner with the entropy paper open in a PDF viewer and a chat panel apologizing that it cannot highlight text in PDF files

I apologize that I cannot directly highlight text in PDF files, as this capability is not available in my toolset. However, I can provide you with a comprehensive and organized list of the most important quotes from Arkady Plotnitsky's paper "The No-Cloning Life: Uniqueness and Complementarity in Quantum and Quantum-like Theories," organized by major themes and with specific section references to make it easy for you to manually locate and highlight them.

Liner

The Liner agent then generates a list of quotes found in the paper with an attached source document that lists these quotes. However, these quotes cite text that cannot be verified (I am unable to click on this text or confirm its placement in the file), and with a bit of CMD+F it becomes clear the quotes read as if they were direct pulls from the source, added to the context window. This is both useless and hurts reputations in professional fields where cited information and accurate evidence attribution are integral and required.

Liner claims it is accurate and trustworthy, but when the generated output isn't 100% correct in its evidence attribution when relaying quotes from published works or generating analysis, I have to go through the files again to constantly verify that my assistant is providing me with correct information. This creates a paradoxical interaction between me and the AI chatbot. I pay for subpar assistance, but I'm left to deal with my problems (which don't get solved), and ultimately, I end up doing more work because the chatbot is less effective than I am at solving or completing my problems and tasks.

Current AI tools are built to learn, not to help. Models learn more from you than you from them, and that is ridiculous.

Let's run the same prompt through Ubik Studio with the same paper. Ubik response:

Ubik Studio with the same paper open, abstract and introduction highlighted in place, and the agent's key highlights listed in the panel on the right

Perfect! I've successfully extracted 8 important quotes from "The No-Cloning Life: Uniqueness and Complementarity in Quantum and Quantum-like Theories" by Plotnitsky.

All notes have been created with precise page references and contextual analysis. Would you like me to retrieve the full text of any of these quotes, or would you like me to create a document synthesizing these key quotes?

Ubik Studio

Ubik agents highlight quotes, allowing users to directly interact with analysis in the PDF viewer, with all annotations saved to the file. These notes aren't just visual; they come with an evidence-rich summary that is verifiable, accurate, and ready to use in your work.

The paper's abstract and introduction with agent-made highlights, and a hover note reading "3 notes here" listing the annotations at that spot
Highlights written into the file itself, with the notes surfacing where they were made.
Ubik's Needs Review panel: seven AI-generated notes awaiting approval, each with page references, an accept and a reject control
Nothing lands without approval: AI-generated notes wait in a review queue.
The All Notes view: 67 notes across the workspace, grouped by the paper they belong to
The workspace after the work: 67 notes, grouped by paper, still there next session.

Working across files, searching academic databases, and delivering traceable, well-cited text in generated output is where Ubik excels.

The way we think about building agents and the user experience is polar from platforms like Liner, ChatGPT, Perplexity, and Gemini. As models get better, don't need the cloud, and become more specialized, the majority of productivity and usability of AI for frontier knowledge work (discovering new science, intense law analysis, publishing to journals, conducting R&D) will be in workflows that require high degrees of human approval, AI trace transparency, and may span long periods of time.

Our approach is markedly different from current cloud-centric stacks, not only in terms of reduced cost, but also in the 95% of workflows where AI integration fails due to:

  • Inaccurate generation and legal issues: incorrect citation or hallucinated work can cause real-world legal trouble and reputation harm.
  • A lack of human approval: if experts and professionals are forced to trust untrustworthy models and agents, more time is spent verifying and correcting, rather than producing and ideating.

Unlike Liner, ChatGPT, Perplexity, or Gemini, Ubik focuses on multi-hop tasks where agents retain pinpoint accuracy while working across many files, searching academic databases, and building expert-level analysis, all stored locally on the device. Here is an example:

A Ubik prompt referencing five papers with @ mentions: review a draft, find supporting quotes in four papers, be critical, search if evidence is missing, then summarize
One prompt, five papers, several dependent tasks: the multi-hop shape benchmarks rarely test.

In this prompt, the Ubik agent has multiple tasks and five separate papers to analyze. Often, these kinds of multi-pronged requests will have a direct impact on the quality and accuracy of the generated output (increased hallucinations). Experts work through complex multi-hop tasks in their high-level research and knowledge workflows. We understand this, and we have built Ubik agents specifically for consistent, accurate, and verifiable output.

Ubik Studio brings usable AI into fields like:

  • Legal teams: track and structure information across diverse source documents for litigation, regulatory research, or case prep. Use the AI agent for brainstorming with citations, or work directly with your organized data.
  • Data and research teams: extract structured datasets from unstructured documents. Turn internal knowledge into alternative data assets with human-in-the-loop review.
  • Consultants and analysts: power your research workflows and build an internal knowledge infrastructure from scattered documents and reports.
  • Scholars and academics: manage sources, search databases, and write with proper citations without leaving your workspace.

Ubik Studio is a local environment where you can upload your documents, and our tools transform them into interactive, searchable knowledge. Explore academic databases (ArXiv, Google Scholar, Semantic Scholar, and more), build structured datasets from unstructured sources, and work with AI that answers with traceable, cited output, all in one place. We grew up with powerful information technology and understand that preference and control build confidence and trust. Ubik gives you control over your data privacy and access to frontier models so that you always work with what suits your needs best.

There is a gap in the literature and research studying multi-hop tasks. There is no standard benchmark for utilizing AI as a co-collaborator in knowledge work over extended periods, or even beyond a chat session. There are numerous benchmarks, such as SimpleQA and HotpotQA, among others. These benchmarks are used to rank and test new AI models in various use cases and field-specific tasks. However, there is an evident lack of successful multi-hop question-answering models or benchmarks explicitly designed to test this process of generation. Although HotpotQA is designed for multi-hop reasoning, it differs significantly from existing multi-hop research tasks.

A HotpotQA example: two highlighted paragraphs about tennis players Brian Gottfried and Peter Fleming above the question "Who has more singles titles" and the answer 21
A HotpotQA item, from the benchmark's own site: context, highlights, question, answer.

This screenshot, taken from the HotpotQA website, gives a random example that shows a question used in the benchmark dataset: "Who has more singles titles, Brian Gottfried or Peter Fleming?"

Above the question (Q), we are shown the context and information the model will use to answer, like a math problem. The highlighted areas are the model-identified sections that contain relevant information based on the words and diction from the question. While these kinds of benchmarks help researchers gain a better understanding of where LLMs struggle in multi-hop reasoning, they do not enable humans to perform high-level multi-hop tasks and knowledge work.

Sometimes we ask the wrong question, and with powerful generative tools, we shouldn't measure efficacy and usability through a lens of replacement and end-to-end solutions. Instead, our measurement procedures should include human entry in the steps between the initial prompt and the golden truth.