AlertLoom: helping a security operations team stop drowning in alerts
A managed security services provider watched dozens of clients' environments and received more alerts than its analysts could read. Real threats hid among thousands of false positives. We built AlertLoom: alert correlation across tools, AI-assisted triage with evidence, playbook automation, case management and client reporting — 24 features that let analysts spend their time on the alerts that matter.

Twelve thousand alerts a day
AlertLoom's SOC lead showed us the queue on a normal day: many thousands of alerts across clients from firewalls, endpoint tools, identity systems and cloud logs. Analysts triaged them one by one, most turned out to be false positives, and experienced staff were burning out on repetitive work.
The real risk was hidden in the noise: an attack rarely shows up as one big alert. It appears as several small, individually unremarkable events across different tools.
“The attack is never one alert. It's six small ones that nobody connected.”
Shadowing analysts
We shadowed analysts across shifts. Their triage followed consistent steps: look up the asset, check whether the user or IP had other recent activity, check threat intelligence, compare with past incidents, then decide. The steps were valuable; doing them manually for every alert was not.
- Manual enrichment: asset, user, IP and history lookups
- Triage of false positives from noisy rules
- Switching between six different tools
- Writing client reports by hand

Weaving alerts into incidents
AlertLoom ingests alerts from every tool into one stream, enriches each automatically, and correlates related alerts — same host, user or indicator within a time window — into incidents. An AI triage assistant summarises each incident, cites the underlying events and past similar cases, and suggests the relevant playbook. Analysts make the call.
Routine responses — isolating a host, disabling an account, blocking an indicator — run as playbooks with an analyst's approval, so response happens in minutes rather than after a chain of emails.
The team
Data engineering for high-volume streams and AI engineering for triage worked side by side with the SOC team.
SOC workshops and shadow-mode evaluation.
Analyst console and client portal.
Streaming ingestion, enrichment and correlation.
Triage summaries, similar-incident retrieval and evaluation.
Console, case management and portal.
Integrations, playbooks and platform hardening.
7 people in total, working as one team.
Decisions we made
Agreed with the SOC lead and the company's security architect.
Replace the SIEM?
- Replace existing tools
- Sit on top of existing SIEM and EDR tools
Our call: Sit on top of existing SIEM and EDR tools. Clients already had tools. AlertLoom adds correlation and triage across them rather than replacing them.
Automatic response?
- Fully automatic containment
- Playbooks run with analyst approval
Our call: Playbooks run with analyst approval. Automatic containment can disrupt business. Approval keeps analysts in control while cutting response time.
Where does client data live?
- One shared cloud
- Regional private deployments
Our call: Regional private deployments. Clients had data residency requirements; regional deployments met them.
How to evaluate AI triage?
- Trust analyst feedback after launch
- Shadow mode against analyst decisions before switching
Our call: Shadow mode against analyst decisions before switching. Shadow mode gave measured evidence of quality before anyone relied on it.
The 24 features
Everything that shipped for analysts, automation and clients.
- 01Multi-tool ingestion
SIEM, EDR, firewall, identity and cloud alerts.
- 02Automatic enrichment
Asset, user, IP and threat intelligence context.
- 03Alert correlation
Related alerts grouped into incidents.
- 04Noise suppression
Known benign patterns suppressed with review.
- 05Deduplication
Duplicate alerts merged.
- 06AI incident summaries
Plain-language summaries citing raw events.
- 07Similar past incidents
Retrieved cases with outcomes.
- 08Incident timeline
Events across tools in order.
- 09Severity scoring
Scored by asset criticality and behaviour.
- 10MITRE ATT&CK mapping
Techniques mapped for each incident.
- 11Playbooks
Isolate host, disable account, block indicator.
- 12Approval workflow
Analyst approval before actions run.
- 13Case management
Notes, evidence and handovers.
- 14Shift handover
Open incidents passed between shifts.
- 15Client portal
Incidents and actions in plain language.
- 16Automated client reports
Monthly reports generated from cases.
- 17SOC metrics
Time to detect, triage and respond.
- 18Multi-tenant isolation
Strict separation between clients.
- 19Regional deployments
Data kept in required regions.
- 20Role-based access
Analysts, leads, clients and admins.
- 21Audit trail
Every action and approval recorded.
- 22Detection rule tuning
Feedback to reduce noisy rules.
- 23Integration health
Alerts when a data source goes quiet.
- 24SSO and MFA
Strong authentication for all users.

Rollout
AlertLoom ran in shadow mode alongside the existing process for several weeks: analysts triaged as usual while we compared AlertLoom's incidents and suggestions with their decisions. Once correlation and triage quality were consistently strong, the team switched over client by client.
- Weeks 1–3Discovery
Analyst shadowing across shifts.
- Weeks 4–5Design
Console and correlation logic prototyped.
- Weeks 6–14Build
Pipelines, correlation, AI triage, playbooks and portal.
- Weeks 15–18Shadow mode
Compared against analyst decisions.
- Weeks 19–20Production
Client-by-client switchover.
What we learned
Correlation beats speed. Grouping related alerts reduced the queue far more than faster triage of individual alerts.
In security, AI must show its evidence. Analysts trusted suggestions only when every claim linked to raw events.
- Next.js
- Python / FastAPI
- Kafka
- ClickHouse
- PostgreSQL
- LLM triage with retrieval over playbooks
- SIEM and EDR integrations
- Private cloud per region

