Buyer Evaluation

    Evaluating Agentic Security Platforms: A Buyer's Checklist

    Every security vendor now claims 'autonomous' or 'agentic' capability. Most of the claims describe a chatbot bolted onto a playbook engine. This checklist gives buyers a structured way to test the claim against evidence before signing a contract.

    2026-05-2712 min readAgentic AIEvaluationProcurement

    Cutting through the marketing claim

    'Agentic' has become a label applied to almost any product with an AI feature. Before comparing vendors on capability, buyers need a working definition of what they are actually testing for. An agentic security platform should be able to perceive an environment, reason about what it observes, decide on a course of action within a defined boundary, and act, then learn from the outcome. Automation that follows a fixed sequence of steps, however sophisticated the branching logic, is not agentic. It is scripted.

    The practical test is simple: present the platform with a situation its rules do not cover, and see what happens. A rebadged playbook engine will fail silently, return a generic answer, or ask a human to write a new rule. A genuinely agentic system will reason from first principles using the telemetry available, propose a course of action, and explain why.

    If a vendor cannot show a live decision trace from an incident type it was never explicitly trained or scripted for, the 'agentic' label should be treated as unproven.

    The seven capability pillars

    A structured evaluation covers seven pillars. Score each one on demonstrated evidence, not on the vendor's own description of the product.

    • Reasoning and model layer: what model performs the analysis, is it a general-purpose LLM API or a purpose-built security model, and how does it handle ambiguity and incomplete data.
    • Agent fabric: how many specialised agents exist, what each one is responsible for, and how they coordinate on a single incident without losing context.
    • Governance envelope: what an agent can do without approval, what requires human sign-off, and what is prohibited outright, and how that policy is enforced technically rather than described in documentation.
    • Data plane: where telemetry lives during reasoning, what is retained, what is sent to third-party model providers, and whether deployment options exist for regulated environments.
    • Integration surface: the specific list of EDR, identity, cloud, SaaS, and ticketing integrations, and how deep each one goes, from read-only visibility to bounded containment actions.
    • Evidence and auditability: whether every agent decision produces a traceable, reviewable record suitable for audit, incident review, and regulatory reporting.
    • Operating model fit: how the platform changes analyst workflow day to day, what skills the SOC needs, and whether the product assumes a staffing model your organisation actually has.

    Proof-of-value scenarios that expose real capability

    Generic demos are designed to succeed. A proof of value should be designed to find the edges of the product. The following scenarios are difficult to fake and reveal whether the claimed capability is real.

    • Novel alert reasoning: introduce an alert type the vendor has not seen in your environment and was not pre-mapped in a playbook. Read the reasoning output as a Tier 3 analyst would review a junior analyst's work.
    • Contested containment: simulate a case where the obvious containment action would break a business-critical process. Watch whether the platform recognises the conflict, escalates appropriately, or blindly executes.
    • Cross-source investigation: trigger an incident that requires correlating identity logs, endpoint telemetry, and cloud activity. A real agent fabric pulls these together into one narrative; a rebadged tool produces three disconnected outputs.
    • Evidence trail quality: after any of the above, pull the full evidence record and check whether a third party, an auditor, a regulator, or a new analyst, could reconstruct the decision without asking the vendor to explain it.

    Red flags to watch for

    Certain patterns recur across products that describe themselves as agentic but are not. Any one of these should slow down a deal.

    • Playbook rebadging: the underlying logic is a fixed decision tree that has simply been renamed 'agent' in the interface and sales material.
    • Opaque model: the vendor cannot or will not describe what model performs the reasoning, where it runs, or how it was trained and evaluated.
    • No policy envelope: autonomy is described as unbounded, with no documented distinction between actions taken automatically, actions requiring approval, and actions that are prohibited.
    • No rollback: contained or remediated actions cannot be reversed cleanly if the agent's decision turns out to be wrong, which turns every automated action into an operational risk.

    Commercial model questions

    Agentic platforms are priced in several different ways, and each model creates different incentives. Buyers should understand which model they are being offered and what it implies for cost at scale.

    • Per-agent pricing: cost scales with the number of specialised agents deployed. Ask what happens to price as you add investigation, containment, or liaison agents over time.
    • Per-ingest pricing: cost scales with data volume. Ask how this behaves during a high-telemetry event such as an incident or a cloud migration, when volume can spike sharply.
    • Per-outcome pricing: cost scales with resolved incidents or actions taken. Ask how 'outcome' is defined, and who arbitrates disputed classifications.
    • In every model, ask for a worked estimate at your current data volume and at twice that volume, so the run-rate assumption is tested before contract signature, not after.

    Procurement and security-review considerations

    Because agentic platforms reason over sensitive telemetry and can take action on production systems, the security review should go beyond a standard SaaS questionnaire.

    • Model provenance: which model performs reasoning, who trained it, and whether it is a general-purpose model called over an API or a model purpose-built and evaluated for security use cases.
    • Data residency: where telemetry is processed and stored during reasoning, and whether local, on-premises, or sovereign deployment is available for regulated data.
    • Privacy: what data is retained after a decision is made, for how long, and whether customer data is used to train models shared across other customers.
    • Third-party dependency: whether the platform's core reasoning depends on an external model provider's uptime and terms of service, and what happens if that dependency changes.

    Operating-model questions

    Capability on paper does not guarantee value in operation. The evaluation should also test whether the platform fits how the SOC actually works.

    • How does analyst workflow change on day one, and what training is required before analysts trust agent output enough to act on it.
    • What is the assumed run-rate model: does the vendor expect the platform to reduce headcount, redeploy analysts to higher-value work, or add a new oversight function.
    • What is the escalation path when an agent is uncertain, and how quickly does a human get looped in.
    • How does the platform behave during a vendor outage or model degradation, and is there a documented fallback to manual operation.

    A scoring rubric for the evaluation

    Score each pillar from 0 to 3: 0 for no evidence, 1 for a described capability with no demonstration, 2 for a demonstrated capability with gaps, 3 for a demonstrated capability with a reviewable evidence trail. A platform scoring below 2 on governance envelope or evidence and auditability should not proceed regardless of its score elsewhere, since those two pillars determine whether the platform can be trusted in a regulated, audited environment.

    How Spharaka answers these pillars

    Spharaka Sphere is built to be evaluated against exactly this checklist rather than around it. SAGE is the purpose-built cybersecurity reasoning model behind the platform's analysis, evaluated specifically on security reasoning rather than treated as a general-purpose language model repurposed for the domain. AuraXP is the multi-agent fabric that coordinates specialised agents for investigation, containment, evidence, and liaison, so complex incidents are handled by agents built for that specific task rather than a single generalist agent stretched across every function. AirWatch is the governance envelope that defines and enforces what each agent can do autonomously, what requires approval, and what is prohibited, with deployment options that include on-premises and sovereign configurations for regulated data.

    The right way to evaluate Spharaka is the same as the right way to evaluate any agentic vendor: run the proof-of-value scenarios above and inspect the evidence trail, not the slide deck.
    Questions

    Frequently asked questions

    How long should an agentic AI proof of value take?

    Four to eight weeks is typical for a meaningful evaluation, long enough to test novel incidents, measure decision quality across multiple scenarios, and observe how the platform behaves under contested or ambiguous conditions. Two-week evaluations test onboarding speed, not reasoning quality.

    Should agentic platforms be compared directly against traditional SOAR?

    No. SOAR executes predefined playbooks; agentic platforms reason and decide. Comparing them on the same axis is a category mistake. The more useful comparison is against other agentic vendors, alongside an assessment of how the platform coexists with the SOAR you already operate.

    Who should own the evaluation inside the buying organisation?

    A joint team works best: detection engineering to assess reasoning quality, SOC operations to test integration reality and workflow fit, the CISO office to review the governance envelope, and infrastructure or legal to assess data residency and commercial terms.

    What is the single fastest way to expose a fake agentic claim?

    Ask for a live decision trace on an incident type the vendor's system has not been explicitly configured for, and ask for the underlying evidence chain behind the conclusion. Rebadged automation cannot produce either convincingly.

    How should commercial terms be tested before signing?

    Request a worked cost estimate at current data volume and at double that volume under the proposed pricing model, whether per-agent, per-ingest, or per-outcome. This exposes cost surprises before they appear in a renewal negotiation.

    Next step

    See it running on your environment

    A walkthrough on your own estate, with your own detections, rather than a canned demo.

    Explore AuraXP™