Extract text from a document
Try local AnyDoc extraction first, then explicitly choose a hosted PDF fallback only when a scanned document needs OCR.
Follow the steps
Upload a supported document into Library and open its item page.
Choose Extract text locally. The job uses local text decoding for plain text and Markdown, or local AnyDoc processing for other supported documents.
If a PDF is scanned and local extraction reports that OCR is needed, choose Scanned PDF options, select a configured OpenRouter or Firecrawl provider, and review the disclosure.
Choose Allow this document only to queue the hosted fallback. The entire PDF is sent to the selected provider only for that job.
Derive text and colors locally
A narrated walkthrough with English captions and a full transcript.
Read the transcript
For an image, choose Extract palette. Rail Cove derives the colors from the image itself, so the same pixels produce the same result.
Select more than one image to extract a combined palette for the collection.
For a supported document, choose Extract text locally. The job stays visible while it runs, and the text is saved with the item when processing finishes.
A scanned PDF can offer hosted OCR separately, only after you approve that document.
What to know
- Hosted processing requires a PDF, explicit allowAI consent, an explicitly selected provider, provider configuration, and a PDF smaller than 8 MB for the hosted worker path.
- A hosted result records its provider and is shown with a reminder to compare the transcription with the original document.
- The app checks visible hosted text and returned reasoning fields for Han characters before saving. This is a conservative character rule that also catches Japanese Kanji and Korean Hanja; it is not a complete language detector.
- The default OpenRouter example uses z-ai/glm-5.3-flash with low reasoning effort and excluded reasoning. Provider/model configuration remains an installation setting.
Troubleshooting
- If local extraction fails without an OCR-needed result, fix the source file or retry after checking the job error; hosted fallback is not an automatic retry.
- If an OCR provider is unavailable, the installation needs its provider configuration. Keep it local prevents a hosted OCR request for the PDF. Smart tagging is separate: an eligible new capture may send bounded extracted text to Mistral AI through OpenRouter. Turn smart tags off in Settings → Notifications before extraction if you do not want that text sent for tagging.
- If the hosted result is rejected for language policy or exceeds the output limit, no partial text is saved and the job reports the failure.
- If the file is too large, split the PDF before requesting hosted OCR. The normal upload limit and the hosted PDF limit are separate.
