A lot of what a document actually says lives in its figures: a clinical algorithm, an org chart, the chart showing the trend the paragraph only gestures at. Docutrain extracts those images during training, has an AI vision model describe them, and shows the relevant ones inside the assistant's answers. Admins curate the result into a browsable gallery, attach videos, and set the cover image.
How images get in
What is extracted depends on the file and the choice you make at upload.
| Source | What is extracted |
|---|---|
| PDF as Document | Individual figures, tables, charts and photos. Graphics under about 5 KB are skipped as decoration, duplicates stored once, and figures re-rendered at higher resolution when that is sharper. |
| PDF as Slide deck | Every page as a full slide image. Decks exported with letterboxing are trimmed to the slide's own edges. |
| PowerPoint (.pptx) | Text only. For slide images, export to PDF and upload as a Slide deck. |
| Admin uploads | Added straight to the image library from the editor. |
That Document versus Slide deck question appears when you select a PDF, and it is the most consequential image decision you make. Reports and handbooks want Document; presentation decks want Slide deck. The full flow is in training your first assistant. Slide extraction then runs in the background after the text is processed — you may see Image gallery is being built in the background, but the assistant already answers questions at that point.
Every extracted image is examined by an AI vision model, which writes a short caption and a longer description. The caption is what readers see; the description is what the assistant searches and reads for context. Editing them is the highest-leverage thing an admin can do here, and edits refresh that image's search index automatically.
Which images end up in an answer
The assistant does not staple every image to every answer. Images are compared against the question, only the ones that stand out are offered up, and the assistant includes only those that genuinely help. Relevance is judged relative to the document rather than against a fixed score, so in a deck where every slide is on-topic the best matches still surface, and in a document where nothing matches no images are attached at all rather than a weak best guess. Section dividers are set aside in favour of content slides, since title cards carry the words of a topic without any of its content. And where several images share a caption, each stays individually reachable.
What a reader sees
A relevant image appears inline with its caption underneath in a styled band, like a figure in a published paper. If the admin turned on Display description in chat, the longer description shows too. Clicking opens the full-screen viewer, and several images in one answer form a temporary gallery you can move through without closing it.
The viewer is the same whether you clicked an image in an answer or opened the document's gallery. A counter at top left shows your position ("3 / 12"); galleries of 15 or more images add a search box filtering by caption, description or file name, reachable with Ctrl+K (Cmd+K on Mac). Zoom runs from 100% to 400% in 50% steps, and clicking the percentage resets. A Downloads menu offers the current image plus any attachments the admin has made available. A thumbnail strip along the bottom hides behind a Show thumbnails button on small screens.
Scroll or click to zoom, pinch and drag on touch, arrows or arrow keys or swipes to move on. While you are zoomed in, swiping is paused until you return to 100%, so panning never accidentally flips the picture.
A document can also offer a curated gallery: a button among its other tools, labelled Images by default with a count badge, opening the viewer in the admin's chosen order. It appears only when the gallery holds at least one image and the Image gallery toggle is on. See the tools beside the chat.
Curating the library
Image management lives in the document editor on the Images tab, which has a Cover Image section and three sub-tabs: Images, Gallery Tool and Gallery Activity.
Two library-wide toggles sit at the top: Include in chat answers, which decides whether images are offered to the assistant at all, and Display description in chat, which decides whether descriptions are shown to readers. With some images on and some off, the toggle reads "Mixed — 12 of 40 on. Click to include all." and resolves to all-on, never all-off.
Under Add Images to Your Document you can upload JPG, PNG, GIF and WebP files up to 10 MB each, 25 at a time. Each needs a Caption before it uploads, because the caption is what makes it findable. Some workspaces also offer a PDF Slides option that turns each page of a PDF into a slide image.
Uploaded Images lists everything attached to the document. Per image you can edit the caption and description in place, toggle Include in chat answers — off means the image stays browsable in the gallery but never enters an answer, which is how a decorative or sensitive picture is handled — toggle Display description in chat, or Replace the file while keeping the caption, description and index intact. Above the grid sit a search box, bulk select and delete, and a Generate All Embeddings button for images not yet searchable.
The Gallery Tool sub-tab controls what readers browse: Gallery Images in display order with drag-to-reorder, Available Images as thumbnails with Preview and Add. Selecting images here does not by itself show the gallery — the Image gallery toggle still has to be on, and emptying the gallery switches it off for you.
Gallery Activity reports who viewed your lightbox images and videos, when and from where, with tiles for Views, Gallery Opens and Unique IPs, a Most viewed rollup, and an Export CSV history. A view counts only after a reader lingers about a second. This is a paid analytics feature.
Covers and video
Every document can have a cover image — the picture on its card in galleries and collections, and at the top of the chat page with the title overlaid. The Cover Image section offers Upload Image, Or generate with AI from a few keywords (usually under 30 seconds), or Or pick from document images. Use dark overlay helps when the light title overlay is hard to read. Covers display at 16:9, ideally 1920 × 1080; off-ratio images are center-cropped, with a warning first. With no cover of its own, a document falls back to the organization's default cover, then to a placeholder tile tinted by its category.
Videos follow a different rule: they are not part of the trained text. On the Videos tab, paste a Vimeo or YouTube link into Video URL and click Add Video; the provider is detected, thumbnails are fetched, and Vimeo supplies the title. Private and unlisted Vimeo links work, and some workspaces allow direct MP4, WebM or MOV uploads. Each video carries a Title, an optional Caption, and an optional Description ("Detailed context used to train the assistant"). Those fields are the video's entire factual content as far as the assistant is concerned: it will point you to a relevant video, but it will not invent claims about what the video shows. A Gallery Tool sub-tab orders which videos readers browse, behind a Video button tied to the Video gallery toggle.
Clicking a video card opens a full-screen player with the platform's own controls and the title, caption and — when enabled — description along the bottom. With several videos you also get arrow-key navigation and a Playlist panel.
Most image complaints trace to two causes: decorative graphics were filtered out or the PDF was uploaded as the wrong type, or the image has a vague caption, no description, or no index. Write captions for the reader and descriptions for the assistant, and the rest mostly takes care of itself.