Inspect each accessible document with an available reader. Establish whether the content is text, a scan or a mix. If a needed page cannot be read, ask for a clearer copy or text and keep that page unresolved.
Record a source document and page or section for each row. For a value combined from several places, retain all contributing references.
Extract the literal value before normalization. Keep the original alongside a normalized date, amount, unit or name where the transformation matters.
Distinguish blank, not applicable, unreadable and absent. Do not turn an unreadable field into zero or infer an identifier because it resembles another row.
When several interpretations are plausible, record the alternatives and the exact field needing confirmation. Use a confidence label with its reason, not an unsupported numerical probability.
Detect repeated documents and duplicate rows by source identifiers and content. Keep duplicate candidates pending an owner decision.
Reconcile document coverage, source line counts, extracted row counts and known totals. Explain omitted headers, subtotals or repeated pages.
Validate destination types, date formats, currencies, units and required fields. Do not combine currencies or convert a unit without a supported conversion.
Return the table with its exception list. If the owner wants a file, use the available writer and inspect the saved content before reporting success.