Picture a lab with twenty years of output on a shared drive: published papers, bench protocols, a couple of theses, and the standard operating procedures nobody can find when they need them. Every new postdoc asks the same six questions in their first month. The answers exist — they are spread across eleven PDFs nobody has read end to end since 2019.
This is a good fit for Docutrain, because the material is already written down and the questions people ask about it are answerable from the text. What follows is the order a group would actually do the setup in.
Getting the PDFs in
From the dashboard, Create Assistant on the Train New Assistant card opens the upload window. The PDF / DOC tab takes PDF, Word, PowerPoint, Excel and Markdown. Give each file an Assistant name — one is suggested from the filename, and you can edit it.
The first decision that matters comes as soon as you select a PDF. Docutrain asks What type of PDF is this? Document is described as "Reports, papers, handbooks — mostly text with some figures," and it pulls out individual figures, tables and charts from the pages so they can appear inline in answers. Slide deck captures every page as a whole slide image instead. For a journal article or a protocol, pick Document. You can change the choice before uploading with the Change button in the File handling row.
The second thing worth knowing is what a scanned PDF costs you. A digital PDF has real text underneath, and extraction is quick. A scan is an image of a page, so Docutrain reads it visually — the slowest and most fragile step in training. It is handled rather than refused: if the service doing that reading stalls or comes back empty, Docutrain hands the same file to an alternative and carries on, with a plain text reader as the final fallback. Protocols photocopied in 1998 will still train; they just take longer. If a paper contains flowcharts or decision trees, an extra pass runs to read them properly, and the processing card says so.
Page count is the one hard limit. Over 300 pages you get a warning that documents this large often exceed training limits. Over 500 the upload is blocked: "This PDF has N pages, which exceeds the 500-page limit for training." The fix is in the message — split the file into parts under 500 pages, upload the first part, then for each remaining part open the document's Retrain tab and choose Add to existing data. The document keeps its link and every setting; only the content grows.
Settings that make a paper behave like a paper
Once training finishes, open the document editor. Four settings do most of the work for research material.
The abstract. The Abstract tab generates an AI-written overview of the whole document from its processed content. Click Generate Abstract, then turn on the Abstract toggle under UI Options → Features → Overview — it stays disabled until an abstract exists. Readers get an Abstract chip in the chat tools that opens the summary with its generation date, a Copy button, and Export as PDF, which emails a formatted copy. Retrain the document later and the abstract is flagged stale.
References. Not the same thing as citations, and the distinction matters here. Citations are generated automatically from the document text and appear under answers. References is a bibliography you write by hand, one entry per line, in the References list field under UI Options → Features → Sources. URLs and DOIs become clickable links automatically, and a live preview shows exactly what readers see. While you are on that sub-tab, check the Citations toggle. With it on, answers carry small Ref chips; selecting one opens a source viewer with the verbatim Source text behind that part of the answer and a chip naming the Page it came from. In multi-document chats it names the document too.
The PubMed link. For a document based on a published article, recording its PubMed identifier adds a PubMed ID button to the document information area of the chat. Clicking it opens a PubMed Article popup with the article's title, authors, journal and year and its abstract, fetched live, plus a View on PubMed link to the full record.
Figures. Everything extracted during processing lands on the Images tab, where each image carries an AI-written caption and a longer description. The caption is what readers see under a figure; the description is what the assistant searches. Both are editable, and editing refreshes that image's search index. To let people browse the figures directly, select and order them on the Gallery Tool sub-tab, then turn on Image gallery under UI Options → Features → Media. The Replace action on an image card swaps in a better scan of a chart while keeping the caption, description and index intact. More in images and video.
Grouping the body of work
One paper is one assistant. A body of work wants a collection. In the dashboard's Collections tab, Create Collection gives you a Collection name, a URL slug, a cover, and a Description — the Generate button drafts that description from the keywords Docutrain extracted from your selected documents.
On the Documents section you tick what belongs. Selected documents move to Selected order, grouped by category, where you can drag them into sequence or use the A–Z and 0–9 quick sorts. That order is exactly what readers see in the collection's sidebar, so a deliberate arrangement — background papers, then methods, then protocols — is worth the two minutes.
Under Chat & navigation, Allow chat across all documents turns a list of papers into a corpus. With it on, the collection page's Chat Mode menu offers Chat across all, and readers can also tick a subset. Citations in a multi-document answer name which document each source came from, so "where does that number come from" stays answerable. Access is set on the collection itself — Public, Passcode protected, or Token based — and one collection passcode unlocks every document inside it, including ones whose own access level is stricter. The editor warns you about that before you save. See collections.
Reading what people actually asked
The part most groups underestimate is the Intelligence tab. It unlocks once a document has received at least 5 questions; until then it shows a progress bar counting toward that.
Overview is plain measurement — questions, conversations, questions per conversation, an estimate of distinct askers, satisfaction from thumbs ratings, and volume over time. Keyword density counts the literal words and phrases people type, and works from the very first question. Once your organization has around 200 questions in total, the Topics panel groups questions into named themes and ranks them, including a Content gaps list ordered by a gap score — higher means answers on that topic more often found no source, got re-asked, or were thumbed down.
Two habits make this useful rather than decorative. Turn on Exclude owner & admin traffic once the document is properly shared; before that, most of the data is your own team testing it. And use the Docutrain Assistant panel — an analyst called Doc, grounded in the question list rather than the document — with its suggested prompt Where are the content gaps? A recurring question your papers do not answer usually means a protocol needs a section written. Document intelligence covers the reports and share links.
If you want to understand why the answers stay tied to your text in the first place, how grounded answers work is the place to start.