Every security vendor now sells an "AI SOC analyst." The demos look identical: an alert arrives, prose appears, a verdict is rendered. Underneath the identical demos sit radically different architectures — some that change your SOC's economics and coverage, some that add a chat window to the tooling you already have. The difference is not visible in a demo. It is visible in the answers to ten questions. Ask all ten, in writing, before shortlisting anyone.
Economics and coverage
1. How does your pricing interact with my data volume? Most SOC cost pathology traces to one design choice: charging for ingestion. If the AI layer sits on an ingestion-priced pipeline, every new log source raises your bill, and you will quietly stop onboarding telemetry — which caps what the AI can ever see. A good answer decouples cost from data volume entirely: pricing by environment size, by alert volume, or by outcome, with an explicit statement that adding a log source costs nothing. Listen for hedges like "we compress before ingesting" — that is the same meter with a discount. The structural alternative is examined in The Economics of Zero-Ingestion.
2. Does the system investigate every alert, or a prioritized subset? Triage exists because human investigation didn't scale. If the AI also triages — sampling, scoring, investigating only "high severity" — you have bought a faster version of the same risk acceptance. A good answer is unqualified: every alert, every source, investigated end-to-end, with the queue's completion rate reported as a metric. Ask what happened to the lowest-severity alert in their last reference deployment.
Trustworthiness of verdicts
3. How are verdicts verified before a human sees them? A single model asserting "benign" is an opinion at machine speed. Mature architectures verify adversarially: a separate agent, with a separate mandate, attempts to break the verdict — to find the evidence path the investigator missed. A good answer names the mechanism, describes what happens when the challenger wins, and reports overturn rates. "The model is very accurate" is not a mechanism.
4. Can I see the evidence behind any conclusion? Every verdict should ship as a case file: the queries run, the raw records retrieved, the timeline constructed, the reasoning chain from evidence to conclusion — reviewable by an analyst or an auditor without redoing the work. A good answer shows you one, unprompted, in the first meeting. If the product summarizes confidently but cannot show its work, you are being asked to trust, not to verify.
The demo shows you the verdict. The architecture determines whether you can afford the coverage, trust the conclusion, and prove it later.
Data handling and safety
5. Where does my raw data live, and who holds it? Some platforms replicate your telemetry into their cloud; others query it where it already lives. The distinction decides your egress costs, your residency posture, and — in regulated sectors — whether the deployment is approvable at all. A good answer is precise about what moves, what persists, for how long, and in whose account. Vague answers about "secure processing" deserve a follow-up in writing.
6. What access does the system actually require? An investigation platform needs to read. It does not need to write, delete, or reconfigure — and any write capability it holds is attack surface an adversary can inherit. A good answer is read-only by default, with scoped, policy-gated response actions as a separately enabled, separately audited stage. Ask for the exact permission grants requested at deployment; a platform confident in its architecture hands you the list immediately.
Fit with what you already run
7. How deep are the integrations — and must my data be normalized first? Architectures that require every source translated into a proprietary schema push a hidden engineering project onto your team, and whatever the parser drops, the AI never sees. A good answer queries sources natively, in their own schemas and query languages, and states integration depth concretely: which APIs, which query capabilities, which sources federated versus synced. Count the sources in your estate their last three deployments actually covered.
8. What happens to my existing SIEM? There are three honest answers — we sit on top of it, we run alongside it, we replace it — and each has consequences for cost, retention, and detection engineering you should hear the vendor reason through. What you should not accept is evasion. If the platform depends on your SIEM, its economics inherit your SIEM's. If it replaces it, ask precisely how compliance retention, historical search, and existing detection content carry over, and in what sequence.
Operations after go-live
9. What do my analysts do once this is running? The right goal is not fewer analysts; it is analysts doing different work — reviewing evidence-backed escalations, hunting, tuning, exercising judgment on the cases that genuinely need it. A good answer describes that workflow concretely: what an escalation looks like, how an analyst challenges or overrides a verdict, and how overrides feed back into the system. A vendor who answers "they just do less" has not run a production SOC.
10. Is your commercial model aligned to my outcomes? The last question ties the other nine together. If the vendor's revenue grows with your data volume, expansion is taxed. If it grows with your alert chaos, they profit from noise. A good answer ties price to something you want more of — coverage, investigated alerts, environment governed — and survives the follow-up: "what happens to my bill if my log volume doubles and my alert quality improves?" The right answer to the first half is nothing.
Scoring what you hear
Run the answers through a simple discipline:
- Require every answer in writing; architecture claims made only verbally have a way of softening by contract time.
- Weight questions 1, 2, 3, and 10 highest — economics, coverage, verification, and alignment are structural and cannot be patched post-purchase.
- Treat any "that's on the roadmap" answer to questions 4, 5, or 6 as a no.
- Validate with a proof of value on your own alerts, scored on investigation coverage, time to verdict, and evidence completeness — the measurement set from MTTI Is the Metric That Matters.
Copilots and platforms both have their place. But they are different purchases with different ceilings, and the label on the deck will not tell you which one is in front of you. The ten answers will.
CAELION publishes this guide because Trace8 was engineered to answer all ten well: zero-ingestion federation, every alert investigated by eight cooperating agents, adversarially verified verdicts with evidence attached, read-only by design. Put the questions to us first — request a briefing.