Tools

    SIEM sizing calculator

    Work out how many events per second your estate produces, how much log data that is per day, and how much storage it needs once retention and compression are applied. Nothing to install, no sign-up, and the numbers stay in your browser.

    Your organisation

    Sets the starting figures. Sources that scale with the estate, endpoint logging and EDR among them, read their unit count from here rather than asking you to type it twice.

    Log sources

    Add a row per source. The same type can appear more than once at different activity levels, which is how you model two chatty firewalls and eight quiet ones. Every rate and event size is editable: replace ours with yours wherever you have measured it.

    01
    devices
    24.06 GB/day
    02
    servers
    15.97 GB/day
    03
    controllers
    9.61 GB/day
    04
    800 endpoints
    5.95 GB/day
    05
    800 endpoints
    25.32 GB/day
    06
    instances
    8.40 GB/day
    07
    gateways
    4.00 GB/day
    08
    accounts
    12.78 GB/day

    Retention and growth

    Hot is the searchable tier your analysts query. Cold is the compressed archive kept for compliance and forensics. They are sized independently, because most deployments keep the archive alongside the searchable tier rather than instead of it.

    Estimate

    2,755Events per second

    Daily ingest, raw
    106.09 GB
    Hot tier, 30 days
    2.18 TB
    Cold archive, 180 days
    7.46 TB
    Total, year one
    9.64 TB
    Year two
    3,031 EPS, 10.60 TB
    Year three
    3,334 EPS, 11.66 TB
    SmallA single node carries this comfortably.
    Where the volume comes from
    • DNS Server28.8%
    • Firewall21.7%
    • Endpoint Detection and Response16.3%
    • Proxy / Secure Web Gateway8.4%
    • Active Directory / Domain Controller7.2%
    • Cloud audit (CloudTrail / Azure Activity)7.2%
    • Windows Event Logs6.7%
    • Email Security Gateway3.6%
    The method

    How the numbers are worked out

    Every SIEM sizing exercise reduces to one multiplication and two divisions. The multiplication is the number of units of a source times the events per second each one produces. The divisions turn that rate into bytes per day and then into the storage each retention tier holds.

    Events per second

    A log source produces events at a rate that depends on what it is and how hard it is working. A firewall carrying a busy internet edge logs far more than one in front of a branch office, which is why each source here has a range rather than a number, and why the same source can appear on several rows at different levels. Total EPS is the sum across every row.

    Daily ingest

    Multiply EPS by 86,400 seconds and by the average size of one event, and you have raw bytes per day. Event size matters more than people expect: a 250 byte DNS query and an 800 byte cloud audit record differ by more than three times for the same event rate, which is why it is editable per source.

    Hot and cold storage

    Raw volume is not what lands on disk. The searchable tier typically compresses to about 70% of raw, and a cold archive to about 40%. Multiply each by its retention period in days and you have the two figures a storage quote is actually built from. They are added rather than blended, because an archive normally runs alongside the searchable tier and not instead of it.

    Growth

    Log volume rises with headcount, with cloud adoption and with every new detection you switch on. The projection applies a compound annual rate to the whole ingest, so year three is the figure worth buying for rather than year one.

    What this cannot tell you

    These are planning estimates. Real event rates move by an order of magnitude with audit policy, with how much of an EDR's telemetry you forward, and with whether you log allowed traffic as well as denied. Use the result to get to the right order of magnitude, then replace any figure you have measured. A week of real ingest from a sample of each source type beats any calculator, including this one.

    Why sizing gets expensive

    The number that actually drives the bill

    Most SIEM licensing is priced on ingest, so every sizing exercise is also a budgeting exercise. Two things follow from that.

    Sizing for the peak, paying for the average

    Event rates are not flat. A working day can run three to five times the overnight floor, and an incident higher again. The rate here is a daily average, which is the right basis for storage. Ingest licensing and node count need headroom above it.

    Volume you pay to store and never search

    A large share of most estates is high-volume, low-signal telemetry that exists for compliance. Sizing it into the searchable tier is the most common way a SIEM bill doubles. That is what the cold tier above is for.

    Spharaka Sphere™ prices on the platform rather than per token, and consolidates SIEM, SOAR, XDR and EDR onto one data foundation, so the same event is not ingested and paid for several times over.

    How AI SIEM differs
    Questions

    Frequently asked questions

    What is EPS in SIEM sizing?

    EPS is events per second: the rate at which your log sources produce records for the SIEM to ingest. It is the primary sizing figure because licensing, node count and storage all derive from it. Multiply EPS by 86,400 and by the average event size to get raw bytes per day.

    How do I calculate daily log volume from EPS?

    Daily raw volume in bytes is EPS multiplied by 86,400 seconds multiplied by the average size of one event. At 5,000 EPS and 450 bytes per event that is about 181 GB of raw log data per day. Event size varies widely by source, so calculate it per source rather than applying one average across the estate.

    How much storage does a SIEM need?

    Storage is daily volume after compression, multiplied by retention. A searchable hot tier typically compresses to about 70 percent of raw and a cold archive to about 40 percent. Size the two separately and add them, because an archive normally runs alongside the searchable tier rather than replacing it.

    How accurate is a SIEM sizing calculator?

    It gets you to the right order of magnitude, which is what a budget conversation needs. It cannot be exact, because real event rates move by an order of magnitude with audit policy, with how much endpoint telemetry you forward and with whether you log allowed traffic as well as denied. Measure a week of real ingest from a sample of each source type before committing to a contract.

    Which log sources produce the most events?

    In most estates the largest contributors are DNS query logging, firewall traffic logs, endpoint telemetry from EDR, and cloud audit trails. Endpoint sources dominate by unit count while network sources dominate by rate per device, so the balance depends on how many endpoints you have relative to your perimeter.

    Should I size for average or peak EPS?

    Both, for different purposes. Storage is sized on the daily average, which is what this calculator reports. Ingest licensing, node count and queue capacity need headroom above the peak, and a working day commonly runs three to five times the overnight floor.

    Does Spharaka charge per event or per token?

    No. Spharaka Sphere is priced on the platform rather than metered per token, so investigations, correlations and autonomous responses can run at scale without a usage-based bill. It also consolidates SIEM, SOAR, XDR and EDR onto one data foundation, so the same event is not ingested and paid for several times over.