Commercial Intent in Human-AI Conversations: A Corpus Audit and Architecture for Website Sales Agents
Information Retrieval
Summary
The gist is being written…
Authors
Benjamin Tannenbaum
Abstract
Conversational sales agents must distinguish questions about products from purchase commitments, preserve explicit requirements, and ground the next action in current business information. We report an aggregate census of 725,219 records in an accessible conversation table provided by Aiso and develop a reference architecture for this setting. All records have distinct non-null conversation hashes. Existing metadata labels identify 41,800 commercial records (5.76%) and 2,387 transactional records (0.33%); their union contains 44,187 records (6.09%). Within the commercial category, 54.31% are labeled English, 41.59% have recorded depth of at least two, and 13.51% have depth of at least four. Commercial-label prevalence varies from 4.87% to 6.35% across three source batches. These measurements motivate explicit separation of corpus inventory, commercial relevance, training eligibility, and observed business outcomes. The proposed architecture combines business-grounded knowledge, provenance-bearing conversation state, and a constrained next-action policy. A quality specification addresses source rights, privacy, label validation, deduplication, and training-test separation. The study is a metadata audit and technical design, not a validation of label accuracy, model training volume, or sales conversion. No raw conversation text or personal identifiers are released.