# Documents

A Space is a data room: folders and documents, plus the facts the extraction
pipeline derives from every processed document. documents.ls lists a Space at a
POSIX-style path; documents.space_contents and documents.folder_contents list
folders and documents; documents.folders_list lists every folder with counts.
documents.get returns one document's metadata and processing status.

Facts: documents.facts_grep searches a Space's extracted facts (BM25 or regex),
documents.facts_cat returns the facts of one document, documents.facts_read
pages through all facts with the current summary and failed or pending
extractions. spaces.summary returns the generated Space summary. Use these to
answer questions about document content instead of downloading files.
documents.content_read (raw OCR/markdown) is gated: call a facts_* operation
for the document first, or pass force=true, or it is rejected.
Known formats (DIAN UBL invoices, Bancolombia and Nequi statements) are parsed
deterministically without OCR or an LLM; documents.jobs_list shows the path in
extractor (for example dian_ubl@1) and in route (deterministic,
parser_failed_fallback or generic), with the parsers considered in
classification. documents.parsers.list shows the known types and what each
one recognizes.

Organising: documents.folder_create, documents.folder_rename, documents.move,
documents.bulk_move, documents.rename. Deleting (documents.delete,
documents.bulk_delete, spaces.clean) is irreversible.

Live documents: documents.live_create makes a collaborative markdown document
or sheet. documents.live_read returns the content and a content_hash;
documents.live_edit changes content (pass the hash back as
expected_content_hash); documents.live_save consolidates and re-extracts facts.
Existing databases: documents.database_read, then documents.database_row_create,
documents.database_row_update ({"status": "done"} moves a task), and
documents.database_row_delete. Person cells such as assignee take a member's
user id or email from spaces.members.

Uploads need the browser: documents.upload_initiate registers metadata,
documents.upload_url mints a PUT URL, the file is uploaded by the client, then
documents.upload_confirm queues extraction. Offer the user the upload area with
navigate instead of uploading yourself. documents.download and
documents.download_bulk mint presigned URLs; never show those URLs in the chat.
They are gated: navigate to the document (or the Space's documents area) first,
or pass force=true, or the call is rejected.

Files in official formats: documents.export converts a document (document_id),
inline markdown, inline html or a table into pdf, docx, xlsx, csv, html or md and
returns a temporary download URL; nothing is stored. Only when the user asks to
keep the file, call documents.export_save, which adds it to a Space as a new
document you can link to. Text exports as pdf/docx/html/md, tables as
xlsx/csv/pdf/docx/html; tables over 2,000 rows only as xlsx or csv.
