If a seller offers you years of inboxes, chats, calls, tickets, pricing records, and operating history, the easiest number to understand is the record count. It is also the number most likely to mislead you.
Your decision is not whether the archive is large. It is whether the archive contains lawful, reconstructable, task-specific evidence that can improve a product enough to repay the acquisition price and the continuing cost of governing it. The reliable way to answer that question is to work backward from a product behavior, validate the data on a representative sample, and expand access only when the evidence clears predefined gates.
The valuable unit is a decision episode, not a file
A model does not become useful merely because it has read more corporate language. The distinctive value in an enterprise archive is the possibility of connecting a situation to the judgment, action, and outcome that followed.
Consider the scope of one proposed corporate-data transaction. Google bid $10 million for Spirit Airlines’ old data, exceeding Mercor’s $7.5 million bid, with the sale still subject to bankruptcy court approval. The proposed package includes approximately 100 million employee emails, 500 million Microsoft Teams messages, more than 30 million customer-service call recordings, pricing and revenue-management data, operational records, and software assets. Passenger profiles and loyalty-program records are excluded, and the remaining data is expected to be de-identified.
The strategic attraction is not just the volume of language. Conversations may reveal what exception an employee noticed, why a team rejected the standard procedure, or which constraint changed a decision. Operational and commercial systems may show what happened afterward. If those records can be linked reliably, the buyer may be able to reconstruct how work was actually performed over time.
But 500 million Teams messages do not equal 500 million useful examples of decision-making. Much of the archive may be routine, duplicated, context-free, or unrelated to the target product. De-identification can also remove the entities and relationships needed to understand why a decision made sense.
My rule is simple: if you cannot describe the unit of learning you expect to extract, you are not ready to value the archive. For many enterprise applications, that unit is a decision episode with six parts:
- The trigger: a customer request, operational event, exception, or business question.
- The contemporaneous context: the policy, account state, inventory, tool output, and other facts available at the time.
- The constraints: permissions, deadlines, costs, dependencies, and rules that limited the available choices.
- The judgment: the option selected and, where recoverable, the reasoning behind it.
- The outcome: what happened next, including escalation, correction, override, customer response, or commercial result.
- The linkage: stable metadata that proves these artifacts belong to the same episode without unnecessarily retaining a person’s identity.
Use that structure to inspect the data. A chat thread without the policy version may be ambiguous. A call transcript without the resulting case status may be unevaluable. A pricing decision without the inventory and demand state may teach a correlation that fails in deployment. The archive becomes valuable when its artifacts can be assembled into examples that match the decisions your product must make.
Write the data thesis before you inspect the sample
Data acquisition is product discovery before it is procurement. Write a one-page data thesis before the seller shows you an impressive demo or lets a large record count anchor the negotiation.
| Decision | What to write down | Evidence required before purchase |
|---|---|---|
| Target behavior | The exact task, user, input, expected action, and unacceptable failure | A representative task set and a baseline from the current system |
| Data contribution | What the acquired corpus contains that your owned, licensed, or generated data does not | Examples showing the missing signal is present and reconstructable |
| Learning mechanism | Whether the data will train model weights, support retrieval, power evaluation, or inform product analysis | A technical path that is compatible with the permitted uses and deletion obligations |
| Unit of value | Decision episodes, resolved cases, linked outcomes, current policies, or another usable unit | A measured yield from raw records to accepted units |
| Success test | The product metric and safety constraints that must improve | A blinded comparison against the baseline and credible alternatives |
| Economic ceiling | Expected incremental value minus acquisition and full lifecycle costs | A conservative business case that does not depend on speculative future uses |
A useful target behavior is narrow enough to test. For example: given a specific service exception, the system should identify the applicable policy, choose an allowed action, and know when to escalate. Improve that sentence until a domain expert can label success or failure without guessing.
Choose the destination before you negotiate the rights
The phrase AI training hides several materially different uses. Separate them before technical design and contract language collapse them into one vague permission:
- Weight training or fine-tuning is appropriate when the archive contains recurring, stable patterns that the model should internalize. It requires strong confidence in provenance, permitted use, retention, and the treatment of derived models because removing influence from model weights is harder than deleting a retrievable record.
- Retrieval is usually a better fit for facts, policies, records, and procedures that change or require provenance at answer time. It preserves access controls and makes correction or deletion more operationally tractable.
- Evaluation data measures whether the system handles realistic work, including exceptions. Keep the final evaluation set separated from training and prompt development, or the result will overstate how well the system generalizes.
- Product intelligence uses the archive to understand workflows, commercial decisions, operational dependencies, and failure patterns without necessarily using the records to change a model. Price this option separately from the model-training case.
The same archive may support more than one route, but that does not mean every route should receive the same data. A policy document may belong in retrieval, a set of adjudicated exceptions may belong in evaluation, and repeated interaction patterns may justify fine-tuning. Architecture is part of data minimization: route each data class only to the mechanism that needs it.
Put every acquisition through four independent gates
A strong business case cannot compensate for unclear rights. Clean rights cannot compensate for unusable records. Treat rights, privacy, semantic quality, and operational readiness as independent gates. Failure at one gate should stop or narrow the acquisition rather than disappear into an average score.
Gate 1: Rights, provenance, and permitted use
Do not accept possession of the data as proof of unrestricted model-training rights. Have qualified counsel examine how each data class was collected, which parties contributed content, what notices and agreements applied, whether third-party material is embedded, which jurisdictions matter, and whether the proposed use is compatible with those conditions. An asset transfer or court-approved sale should not replace record-level diligence.
Build a rights matrix with one row for every meaningful data class. Include its originating system, data subjects, contributing parties, date range, geography, proposed AI use, allowed transformations, retention conditions, deletion requirements, and evidence supporting the seller’s authority. Separate employee communications, customer interactions, operational telemetry, commercial records, attachments, and software assets; they do not necessarily carry the same rights.
Stop if the seller cannot produce a credible inventory, relies solely on corporate ownership as its rights argument, or cannot separate disputed and excluded data. Those gaps can leave you with an archive that is technically accessible but commercially unusable. Contract structure must be reviewed for the transaction and relevant jurisdictions; a diligence checklist is not a substitute for legal advice.
Gate 2: Privacy preservation without destroying the signal
De-identification is a transformation to test, not a label to trust. Free-form email, chat, audio, attachments, and metadata can contain identity clues in places a simple name-removal pass will miss. Conversely, aggressive removal may erase account continuity, reporting relationships, chronology, locations, or tool states that make an episode understandable.
The proposed Spirit package is instructive because it combines de-identification with explicit exclusion of passenger profiles and loyalty-program records. Scope exclusion is often more reliable than trying to transform every available field. If a data class adds little to the target behavior, do not acquire it.
Test the privacy pipeline on representative records before transfer. Check the text, audio, metadata, attachments, and join tables. Attempt to recover identities using the transformed fields available to an internal user or model. At the same time, measure whether consistent surrogate identifiers, timestamps, policy versions, and outcome links still allow reconstruction of the intended episodes. Set the acceptable privacy leakage and semantic-loss criteria before seeing the results.
Gate 3: Semantic quality and historical relevance
File integrity tells you whether a record opens. Semantic quality tells you whether it can support the task. Sample across time periods, systems, business processes, locations, and outcome types. Combine a broad sample with targeted sampling of rare exceptions because a broad sample alone can be dominated by routine work.
Track a yield funnel from raw records to rights-eligible records, readable records, reconstructable episodes, adjudicable examples, and examples that survive privacy transformation. The denominator matters at every stage. Reporting only the final number of promising examples conceals how much of the purchased corpus will become storage and governance overhead.
Historical depth can reveal how a company adapted, but age is not automatically an advantage. Policies, tools, vocabulary, customer expectations, and operating constraints change. Preserve time and policy version as features, then test older periods separately. If a model learns a superseded workflow as though it were current policy, more history can reduce product quality.
Gate 4: Security and operational readiness
The purchase price is only the entry cost. Inventory the work needed to extract proprietary formats, recover thread structure, link outcomes, transcribe recordings, handle attachments, remove duplicates, transform sensitive fields, label episodes, store the corpus, control access, monitor use, respond to incidents, and execute deletion requirements.
Require an operational plan for where the data will be processed, which people and vendors can access it, which environments are allowed, how derived datasets and embeddings will be tracked, and how deletion will propagate. If the organization cannot identify every downstream copy and derivative, it cannot make a credible retention or deletion commitment.
Validate on a sample and price the demonstrated yield
Do not make the full acquisition the experiment. Use staged access to determine whether the archive clears the four gates and improves the target product before taking custody of the whole corpus.
- Freeze the product task, baseline, evaluation set, safety constraints, and rejection criteria before examining the sample.
- Define a sampling plan across time, channels, workflows, and outcomes. Require the seller to disclose how the sample was selected so that a curated showcase is not mistaken for representative yield.
- Inspect and process the sample in a controlled environment with the minimum access necessary. Do not pull the complete archive into your normal development stack for convenience.
- Run the rights and privacy transformations first. Examples that cannot be used lawfully and safely should not influence the quality estimate.
- Construct decision episodes and have domain experts adjudicate whether the context, action, and outcome are sufficient. Record disagreements rather than forcing ambiguous examples into a clean label.
- Compare the acquired-data approach with realistic alternatives: your existing data, retrieval over approved knowledge, tool integration, synthetic examples, or a narrower licensed dataset.
- Measure both product lift and failure movement. An average improvement can hide worse performance on rare, costly, or sensitive cases.
- Calculate the accepted-unit yield and full lifecycle cost, then expand only if the results clear the thresholds committed in advance.
Use a conservative economic ceiling: expected incremental product value minus the acquisition price, extraction and transformation work, expert labeling, security, legal review, storage, evaluation, serving, monitoring, retention, and eventual deletion costs. Keep speculative uses outside the base case. If commercial analytics or future product options matter, value them explicitly rather than letting them silently inflate the training-data bid.
Raw record count is a poor payment unit because it rewards volume whether or not the records survive diligence. Negotiate around accepted partitions, measurable usable yield, or staged rights where practical. Any representations, warranties, indemnities, audit rights, acceptance remedies, or contingent payments need transaction-specific legal drafting.
Make the contract match the architecture
Hand counsel and procurement a technical schedule, not just the phrase AI training. The schedule should identify:
- The exact systems, date ranges, formats, metadata, attachments, and exclusions being transferred.
- The allowed products, model-development activities, evaluation uses, analytics uses, users, affiliates, and vendors.
- The treatment of cleaned records, transcripts, labels, embeddings, evaluation sets, model weights, and other derived artifacts.
- The access, security, location, retention, deletion, incident-response, audit, and downstream-transfer controls.
- The acceptance tests for schema completeness, readability, linkage, privacy transformation, sample representativeness, and rights documentation.
- The obligations that survive the end of the license or the seller’s continuing involvement.
Put the walk-away conditions in the investment memo before negotiations begin. Walk away or narrow the scope if rights cannot be audited, privacy controls destroy the necessary context, usable yield falls below the agreed floor, the data produces no material lift over safer alternatives, or lifecycle costs erase the expected value. A high bid from another buyer does not resolve any of those failures for your product.
Key takeaways
- Buy reconstructable decision episodes, not impressive quantities of messages, recordings, or files.
- Define the target behavior, learning mechanism, evaluation, economic ceiling, and rejection gates before inspecting seller-selected examples.
- Evaluate rights, privacy, semantic quality, and operational readiness independently. One strong dimension cannot cancel a failed gate.
- Measure yield after rights filtering and privacy transformation, not before them.
- Separate training, retrieval, evaluation, and product-intelligence uses because they require different data, controls, economics, and permissions.
- Expand access in stages and make the contract describe the actual technical architecture and derived artifacts.
Your next step should not be a bid or a broad request for all available data. Write the one-page data thesis, define the walk-away gates, and ask for a representative sample plus a data-rights matrix. If the sample cannot demonstrate both lawful usability and measurable product lift, stop. If it can, expand deliberately. That turns a contest for a large archive into a defensible AI product investment.
References








