This paper introduces agentic data cracking, a method that dynamically structures unstructured data as a byproduct of agent reasoning: when an LLM agent opens a document to answer a query, a secondary sub-agent extracts structured information likely useful for future related questions. On the FanOutQA benchmark the technique cuts cost by 53% while maintaining accuracy, improving token efficiency for enterprise applications that reason over documents.
