# Going further
# Background autopilot runs
Autopilot can build journeys without you watching. When you create journey structures from a requirement, the confirmation offers Start autopilot, which queues each new journey as a background run; the requirement's Journeys tab carries the same Start Autopilot button for journeys created earlier. Each run builds one journey. Several run at once, up to the cap for your organisation, and the rest wait their turn. When a run finishes it posts an Autopilot notification with its result, for example "4 completed, 0 failed out of 4 checkpoints".
A background run goes through the same stages as a run you watch: it plans from the structure, executes against your application in a live browser, writes standard Virtuoso steps into the journey, and finishes with the AI-free validation execution. The difference is when you get involved. Rather than steering as it goes, you review the result when the notification arrives: open the journey, inspect the generated steps and the validation verdict, and start Autopilot on it again if it stopped short or the re-run didn't pass. Part 5 of the step-by-step guide describes what to look for from step 9 onwards.
Background runs suit a batch of similar journeys once you trust the structure and your Autopilot directives. For the first journey in a new area, an attended run is the better choice: that is where your answers to Autopilot's questions, and your adjustments to its plan, make the most difference.
# Keeping tests up to date
Touchstone is designed as a maintainable loop rather than a one-shot generator. Change propagation is enabled per project and is being rolled out progressively: where it isn't enabled, nothing in this section applies and generation behaves exactly as described above.
TIP
Part 8 of the step-by-step guide walks the review through; this section is the model behind it.
When a source document changes, upload the new version over the existing source rather than adding it as a new one: replacing in place is what preserves the link to the requirements already derived from it. That produces a new version of the document and a comparison against the previous one. The change is classified by how much it actually matters, and only a substantive change raises anything: a reformat or a typo fix is recognised as cosmetic and passes without disturbing your tests. A project is only flagged if at least one of its requirements is actually linked to that source, so a shared document changing doesn't create noise in projects that don't depend on it.
Where the change does matter, the requirements linked to that source are flagged for review. Opening a review conversation loads each affected requirement together with everything that has changed in the source since that requirement was last reconciled, so a requirement that sat through three revisions is reviewed against all three at once, not just the most recent.
Each affected requirement then gets one of three verdicts:
| Verdict | Meaning |
|---|---|
| Mandatory | The change contradicts the requirement: a corrected version is proposed |
| Optional | The change adds something the requirement could cover: a refinement is proposed |
| No change | Reviewed and still correct: returned unchanged with a one-line explanation |
In the proposals list, a requirement with a Mandatory or Optional change is marked To approve and waits for you; a requirement that still holds is marked No change and is approved on your behalf, so a review of seven requirements where two changed is two decisions, not seven.
No change is a real answer, not a skipped one. Every flagged requirement gets an explicit verdict, and a no change verdict still has to be saved: that is what records the review as having happened. The flag clears because the review happened, not because something was edited, and a flag that persists means something linked to the source is still unreviewed.
Propagation is one hop at a time, and always through you. Saving a changed requirement flags the journeys linked to it for review. It doesn't rewrite them. Those journeys are then reviewed against the current version of the requirement in their own conversation, with the same three verdicts. Nothing moves without a decision from you at each stage, and flags clear themselves once nothing downstream is lagging.
When a journey fails, Autopilot classifies the failure, proposes a fix on the journey, and re-runs it to confirm. Part 6 walks that through.
Because requirements, data tables and environments are versioned, the state of your test estate at any point is reconstructable, and changes are reviewable rather than silent.
# Getting the best results
Output quality tracks the quality and focus of the context you provide:
- Select sources deliberately. Upload a well-structured Knowledge Base, then choose only the sources relevant to the feature under test for each generation. Supplying everything at once dilutes retrieval and degrades results.
- Balance your sources. Generation weights towards whatever it has the most material on: extensive API documentation alongside thin UI documentation will skew requirements towards the API. Curate accordingly.
- Be specific in the Chat. State the feature, the behaviours you care about, and the test cases you have in mind. Precise prompts outperform broad ones; overloading context reduces accuracy.
- Encode standards once. Put conventions and generation rules in Directives; keep the underlying policy and regulatory detail in the Knowledge Base.
- Treat generated output as a first pass. The generated set is a structured starting point: the value of the Chat is applying your judgement about what's missing or superfluous.
Once you reach the journey-building stage, a few more apply:
- Seed, then scale. Build the first journey in an area with Autopilot, extract the stable parts as library checkpoints, then generate the remaining similar journeys against those shared building blocks: faster, and more consistent, than generating every journey from nothing.
- Tidy up in bulk. A batch of newly generated journeys usually needs the same housekeeping applied to all of them. Select several in the goal's journey list and set assignee, publish state or lifecycle status across all of them at once, rather than opening each in turn. Generated journeys arrive as drafts, so publishing the batch you're happy with is a single action.
- Prefer environments for varying inputs. Data tables are supported, but environment variables are the better mechanism for URLs, credentials and anything that changes between environments. Where a structure does carry a data table, its checkpoints reference columns by name rather than by literal value, so the generated steps stay parameterised.
- Well-engineered front ends yield the best results. Accessible, standards-following applications are the strong case. Elements inside iframes or shadow DOM, and unconventionally built pages, can require manual adjustment of the generated steps.
- Structure first. Autopilot doesn't introduce loops or conditional branches. Conditional steps already present in the journey are expected to execute; journeys containing loops aren't yet supported in Autopilot runs.
# What to expect from an Autopilot session
- One session per journey. Each Autopilot session is bound to a single journey, and the number that can run at once is capped.
- Authoring is slower than playback. While generating steps, each interaction includes model reasoning time, so an authoring run takes longer than the resulting journey will. The validation execution and all subsequent runs execute at normal speed.
- The browser is provisioned in the background while you read and adjust the plan, so there is nothing to wait for before planning.
- It runs on the standard browser and shared infrastructure. Fixed IP addresses and custom certificates aren't supported; cross-browser verification is done by normal execution after the journey is built.
# Current limitations
Constraints to be aware of:
- Images inside documents are not read. Only the text of an uploaded file is extracted. Screenshots, diagrams and annotated images embedded in a PDF or Word document are skipped, which matters most for user guides whose meaning lives in the pictures. Upload those images as sources in their own right, where they are analysed.
- Scanned documents can't be read. A PDF whose pages are images rather than text, a scan or a photocopy, has nothing for Touchstone to extract, so the upload is rejected rather than partially processed. Convert it with any OCR tool and upload the searchable version. This is a current boundary rather than a permanent one.
- Very large sources can fail while summarising. Ingestion of large documents is still being hardened. If a large file is rejected, split it into sections and upload those.
- Generation is not deterministic. The same sources and the same prompt can produce different output on two runs: a different number of acceptance criteria, a different split between requirements, a different route to the same coverage. The same is true of Autopilot: the same journey against the same application can produce a different number of assertions, different waits, or a different route to the same objective. This is inherent to the models behind every generative feature here. Treat one run as a starting point to refine, not as a fixed answer to reproduce.
- Retrieval is selective. With a very large Knowledge Base, guide the conversation towards the sources that matter; extremely large individual documents are ingested via their summaries, with some loss of detail.
- Template directives apply to requirements only. Journey generation can be steered with instructions and conventions, but its output structure isn't template-configurable.
- Directives are guidance, not a guarantee. They shape generation reliably in practice, but they're applied by an analysis step, not enforced by a rule engine: expect them to be followed, but verify on output that matters.
- Touchstone is not, by itself, a compliance solution. To test against a standard, capture the applicable rules in your Knowledge Base, encode the structural requirements as template directives, and direct generation at them. A template guarantees the section exists on every requirement; it doesn't guarantee the content in it satisfies the regulation, and nothing here certifies conformance. Where you need auditable enforcement rather than best-effort steering, treat the generated output as the input to your own review gate.
- Asset reuse is uneven. Existing goals and environments are reused. Library checkpoint reuse is still being introduced, so generation creates new checkpoints instead. Other asset types, such as scripts and extensions, aren't included.
- Approving and rejecting are interface actions. You can shape a proposal through the Chat, but you can't ask Touchstone to approve or reject one: that decision stays with you deliberately.
- A created asset is closed to the conversation. Once created, a requirement or journey can't be refined further through the Chat. Duplicating one to iterate on it is planned but not yet available.
- Change review is rolling out. Sources are versioned and requirements record the version they were derived from. The review flow that compares a requirement against everything that has changed in its source is not yet available in every environment.
- Knowledge Base filtering is being reworked. The filters for organisation-level, project-level and archived sources will change in a future release.
Specific to journey building and repair:
- Generated wait durations track the authoring session, not your application. See Part 5, step 12.
- Some control types are unreliable. Date and time pickers, range sliders, random data generation and final-message validation are the known weak spots. A range slider that can't be set to an exact value is the common case: Autopilot pauses and asks how to proceed rather than writing the wrong number, but you'll need to answer.
- File upload can fail even when the file is present in the environment; a manual step adjustment is sometimes needed.
- The preview panel is view-only. It shows the session, but interactive step editing remains in Live Authoring.
- Journeys containing loops aren't yet supported in Autopilot runs; fixed IPs and custom certificates aren't available on Autopilot's browser infrastructure.
← Step by step Overview →