# Knowledge Base

The first stage of the pipeline. This page explains what the Knowledge Base does with your sources.

TIP

Part 1 of the step-by-step guide adds sources and checks them.

The Knowledge Base holds the source material Touchstone reasons over. It accepts uploaded files such as system overviews, functional specifications, and domain and regulatory material, and operates at two levels:

  • Organisation level. Cross-cutting material relevant to every project, such as company-wide policies or compliance documentation. Managed by the organisation owner, and can be assigned to specific projects.
  • Project level. Sources specific to a single project or application under test.

A wide range of formats is supported:

  • Text. Plain text, Markdown, CSV, JSON and JSONL.
  • Office documents. PDF, DOC and DOCX, PPTX, XLS and XLSX.
  • Images. PNG, JPG and SVG, interpreted by a vision model so diagrams and annotated screenshots contribute to analysis. This applies to an image uploaded as a source in its own right. An image embedded inside another document is not read: see below.
  • XML formats. Including BPMN and draw.io diagrams, WSDL, XSD, XMI and UML schemas, and test reports such as JUnit, TRX, NUnit, Cobertura, JaCoCo, Checkstyle, PMD and SpotBugs. Each format is interpreted according to its own conventions rather than as generic markup.

Only text is extracted from documents. A format-specific loader reads each item's content on upload and indexes it for retrieval, and for PDFs, Word documents and the rest, what it reads is the text. Images embedded in those files are skipped entirely, so an annotated screenshot inside a user guide contributes nothing. This is a deliberate boundary: a PDF can run to hundreds of pages, and analysing every embedded image is expensive enough to need handling carefully. The route round it is direct: export the images that matter and upload them as their own sources, where they are analysed, describing what each one shows.

Each item shows its progress: uploading, then summarising while its content is analysed, then it carries a NEW badge (UPDATED if you replaced the file on an existing source) once it can be used in the Chat. An item that can't be processed is marked failed and kept so you can inspect it. A file with no text to extract at all, a scanned PDF, for example, can't be read and is rejected on upload; run it through OCR first and upload the searchable version it produces.

Summary first, full document when it's needed. Once a source is ready, Touchstone works from a mix of the two. The generated summary is the compact representation it searches and reasons over by default; where that isn't enough to answer properly, it fetches the whole document and reads the original content. That decision is its own, made as it works, so detail that only exists in the source isn't lost. A title, description and summary are all produced automatically, and you can edit any of them. The summary is generated as a structured breakdown rather than free prose, classification, content scope, entities, operations, business rules, actors and test considerations, so it's quick to scan for anything that's been misread. Saving a revised summary re-indexes the source, and the source is briefly unavailable while that happens, so the correction takes effect from the next generation onwards. An inaccurate summary is still quietly expensive, because the summary is what gets searched: a source that's been misread may not surface at all for a request it should have answered. Read the generated summaries once after upload. It's the highest-value five minutes in the whole setup.

Organising sources. Sources can be tagged at project level to make them easier to find and filter; tags don't exist at organisation level, so organisation-wide material isn't tagged. Sources can also be archived, which takes them out of use without losing them, and archived items can be restored or deleted outright.

Which projects a source reaches. Open a source and its Projects tab lists the projects it's assigned to. The default is worth knowing, because it's permissive rather than restrictive: a source with no projects selected is available to every project. Selecting projects narrows it to those; leaving the selection empty opens it to all of them. If a source is sensitive to one team, assign it deliberately rather than assuming an unassigned document stays put.

At generation time, Touchstone retrieves the most relevant material for the request rather than loading everything into context. When generating requirements it can also go back for more mid-answer: if the analysis reaches a gap, it searches the Knowledge Base again or loads a specific document in full, without waiting to be asked. Those follow-up lookups are capped per response, to keep the context focused and the cost predictable. Journey generation works differently: it fetches documents on demand as it writes, and can also expand a project asset (a journey, library checkpoint, goal, data table or environment variable) to its full shape when it needs the detail.

This is why source selection matters: a curated, well-categorised Knowledge Base with deliberately selected sources produces measurably better output than an exhaustive one (see Getting the best results).

Last Updated: 9/30/2026, 2:07:07 PM