Bicycle.LLC

← The work

CanonicAI is software for corpus owners and builders that turns books and papers into cited data.

CanonicAI — corpus in, canonical data out; the public face of the engine at canonicai.com.

I built CanonicAI for people who need to turn a body of books and papers into data they can inspect and use. The first version covered too many kinds of work: books, research, jobs, compensation, segmentation and business material. Narrowing it to canonical data made the product clearer and the engineering better. The difficult part was getting the same multi-step extraction to behave consistently across thousands of runs, then keeping the source, identity and cost attached to the output. A prompt can change overnight. Records that other products rely on have to survive reruns, duplicate language and authors who disagree. CanonicAI therefore keeps the books and papers visible in the result. A reader can see which source supports a relationship, and a builder can use that same record in software. It now supplies the cited data behind the products I build. The longer aim is a public canon where a person or an AI agent can inspect a field's ideas, evidence and disagreements before using them.

Who it is for

CanonicAI is for anyone who owns or can provide a body of books and papers, and for builders who need to use that material in software or an AI agent. It turns the corpus into records they can inspect and query, with sources visible instead of buried in PDFs.

The problem

Books and papers are useful source material, but a chat model over PDFs can produce answers that are difficult to defend. Citations drift, one concept appears under several names, and sources that disagree are blended into one voice. Manual extraction adds another problem: the same multi-step work must be repeated and checked for every document. The result is often slow, inconsistent and hard to reuse in a product, guide or agent.

What I built

CanonicAI is software for corpus owners and builders that turns books and papers into structured, queryable data with a source on every claim. For books, it follows real chapter boundaries; for papers, it extracts research concepts, measurement instruments, citations and findings. A repeatable production process reconciles different names for the same idea, gives each concept a stable identity, records typed relationships and preserves genuine disagreements. It can rerun the same work without creating duplicate outputs, while retaining the source trail and production cost. People can browse the resulting records, and software or an AI agent can call them. The public canon shows the system at work with 155 domain models, 3,716 stable construct IDs and 8,899 cited relationships.

What is new in it

  • The production process can apply the same multi-step extraction across thousands of documents, avoid duplicate outputs on reruns, retain the source trail and track cost.
  • Different authors can use different names for the same idea while CanonicAI maintains one stable identity and preserves their original terms as aliases.
  • Relationships retain their supporting sources, and disagreements remain part of the data. Readers and builders can inspect the evidence before reusing a connection.
  • Books are divided at their real chapter boundaries, keeping extracted claims connected to the structure in which each author developed the argument.

Where it stands

CanonicAI is becoming a shared source of cited field data that people can inspect and software can reuse, so useful knowledge does not disappear inside a chat response. The direction is already visible at canonicai.com, where 155 domain models, 3,716 stable construct IDs and 8,899 cited relationships are available to browse or call.

More screens