• Artificial Intelligence
  • Cybersecurity
  • Monitoring

The alerts nobody reads

Why SMB monitoring breaks down before the infrastructure does, and how an AI-assisted analysis pipeline can give teams back what they’re actually missing: consistency.

In most infrastructures we come across, the problem isn’t a lack of signals. It’s the opposite: the signals are all there, correctly logged, sitting exactly where they should be — and nobody reads them. Not out of negligence, but because actually reading all of them, every day, isn’t a job a person can do.

The noise has already won

The numbers on so-called alert fatigue are by now fairly uncontroversial. Organizations receive on average around 3,000 security alerts a day and leave nearly two-thirds of them unhandled; between 20% and 30% are never investigated in time. The 2025 SANS Detection and Response Survey reports that 73% of security teams name false positives as the main obstacle to detection, and Microsoft and Omdia’s State of the SOC 2026 report estimates that 46% of alerts turn out to be false positives: almost half an analyst’s workload produces no security value at all.

The cost isn’t just the wasted time. It’s the desensitization. A system that’s always shouting stops being an alarm system and becomes background noise: 76% of organizations name alert fatigue as one of their main operational concerns, and roughly three-quarters of analysts say they no longer have time for proactive work. This is the point where monitoring has formally broken down while continuing to technically run.

SANS 2025 Detection & Response Survey, analysis by Stamus Networks — stamus-networks.com

Undersized teams, poorly tuned systems

In small and medium businesses, this dynamic compounds with two structural conditions.

The first is staffing. The 2025 ISC2 Cybersecurity Workforce Study, covering over 16,000 professionals, finds that 95% report at least one skills gap in their organization and 59% call it critical or significant; 88% have suffered at least one concrete security consequence tied to that gap. A third of respondents simply say they don’t have the resources to properly cover their own perimeter. In an Italian SMB this often translates into one person — sometimes half a person — covering everything: network, backups, workstations, telephony, and, on top of that, security.

ISC2 Cybersecurity Workforce Study 2025 — isc2.org

The second is configuration. Monitoring systems, where they exist at all, are almost always installed and poorly tuned. The symptoms repeat with suspicious regularity:

  • thresholds left at the vendor’s default values, calibrated for an average infrastructure that isn’t the customer’s;
  • no baseline: there’s no definition of “normal” for that environment, so there can be no definition of “abnormal” either;
  • alerts on everything that’s technically monitorable, instead of on what actually impacts the service;
  • notifications that all funnel into a shared inbox nobody owns, with no ownership or escalation path;
  • no correlation across sources: storage, hypervisor and application each tell a third of the same story in three different places;
  • log retention too short to notice a degradation that develops over weeks.

The result is a setup that generates work instead of reducing it, and that misses exactly the events it was supposed to catch: the slow ones.

The human limit isn’t skill. It’s consistency.

It’s worth being precise about where the real advantage of an automated system actually lies, because it isn’t where it’s usually described.

An experienced sysadmin is far better than any model at interpreting a fault that has already manifested: they have the context, they know the infrastructure, they know what was touched last week. The real point is different. The human eye is excellent at spotting the macroscopic anomaly and structurally blind to gradual change. An event that goes from three to thirty occurrences a day over the course of a week doesn’t “look” different when you scroll through a log: it looks normal, because you saw it yesterday too. The difference only becomes visible by counting — and counting continuously, on the same metric, for weeks.

This is exactly the kind of work a well-calibrated pipeline does without degrading: it doesn’t get tired on Friday afternoon, it doesn’t skip a check because there’s an emergency somewhere else, it doesn’t lower its attention threshold because “that error has always been there.” Consistency, not intelligence, is the function actually being delegated.

A real case: the disk that had already failed

A customer reported a symptom as annoying as it was vague: daily operational stalls lasting a few minutes, after which everything went back to working normally. No errors reported to users, no service visibly down, no alerts. The infrastructure was a Dell server acting as a Hyper-V host.

Hardware monitoring showed nothing. iDRAC hadn’t raised any predictive failure: disk health parameters were within the vendor’s expected thresholds. From the controller’s point of view, the infrastructure was healthy.

We exported the host’s Event Viewer logs in full and fed them to a language model along with the necessary context: the machine’s role, the nature of the reported symptom, the time window, the storage configuration. The analysis isolated a pattern that wasn’t an error: a disk sector was being rewritten repeatedly, and the frequency of that behavior had increased roughly tenfold over the previous week.

None of the individual events making up that pattern was, by itself, an alert. The signal wasn’t in the event: it was in its derivative.

That data was enough to build a detailed technical report for the manufacturer, who authorized the disk replacement before the predictive failure threshold was ever crossed. Once the disk was replaced, the operational stalls stopped: those few-minute pauses had been repeated write attempts on the degraded sector, which were tying up the storage underneath the entire virtualized host.

The interesting part isn’t that “AI found the fault.” It’s that the fault had been written in plain sight in the logs for weeks, accessible to anyone who had counted those events day by day. Simply, nobody was doing it — and no system had been configured to do it either.

How a pipeline like this is built

An isolated case is an anecdote. What turns it into a method is industrializing it into a repeatable flow. The architecture we propose is built on five stages.

The five stages of AI-assisted observability

Deterministic where it’s enough, a model where it’s needed

The most common mistake is handing everything over to the language model. That’s inefficient and expensive: counting occurrences, calculating deviations and applying dynamic thresholds is work for a query and a few lines of statistics — deterministic and verifiable. The model comes in afterward, on a volume of data already reduced by orders of magnitude, and it’s used for what it does better than a rule: reading heterogeneous sources together, forming a hypothesis and explaining it in language a technician can verify in ten minutes.

The guardrails aren’t optional

  • No destructive or irreversible action is ever executed automatically. Reboots, host isolation, deletions and configuration changes remain human decisions.
  • Every model output must come with the raw evidence it’s based on, independently verifiable by the technician. A hypothesis with no log lines to back it up isn’t usable.
  • Log handling has to be designed around the nature of the data: minimization, pseudonymization where necessary, and a deliberate choice between on-premise processing and external services.
  • The pipeline has to be measured like any other detection system: false positives, but above all false negatives. A system that never gets it wrong is almost always tuned far too loosely.
  • Calibration isn’t a project activity, it’s ongoing maintenance. The infrastructure changes, and so does the baseline.

”But this is what a SIEM does”

Yes. Most of this workload is exactly what enterprise SIEMs exist for, and nobody here is arguing otherwise. The problem is real-world accessibility, not technical capability.

A SIEM isn’t a product you install, it’s one you adopt. It requires licensing proportional to the volume of ingested data, a source-onboarding project, a tuning phase that lasts months and, above all, stable staffing to actually read what it produces. Without that staffing, a SIEM becomes the company’s most expensive alert generator: it adds noise to an organization that was already drowning in it.

And it’s precisely the segment that can’t afford one that’s most exposed. The Clusit 2026 Report records 5,265 cyber incidents globally in 2025, a 48.7% increase over the previous year — the highest ever recorded. It places Italy at 9.6% of worldwide attacks. In Italy, SMBs make up roughly 72% of targets, with nearly one in four having suffered a breach in the last three years and an average cyber maturity score stuck at 55 out of 100, below the passing threshold. Internationally, the Verizon 2025 DBIR notes that ransomware extortion appears in 88% of breaches involving small and medium businesses, versus 39% for large organizations.

Clusit 2026 Report, as covered by Innovation Post — innovationpost.it

Verizon Data Breach Investigations Report — verizon.com. The figure cited is from the 2025 edition; the page now hosts the 2026 edition.

Pipelines don’t replace a SIEM. They’re the step that precedes it for those who can’t yet afford one, and the multiplier that cuts the noise for those who already have one. In the right order: first put someone or something in place that can actually read the signals, then increase the number of signals collected. Doing it the other way around is the most reliable way to waste a security budget.

Where we fit in

Fenix & Vega works exactly in this gap: between what a technology promises and what the customer’s organization can actually absorb today. We don’t propose the most advanced tool on the market, but the one the customer will still be able to keep alive in twelve months.

Depending on the case, this means we design and run the pipeline as a service ourselves; or we build it together with the in-house IT team and train the people who’ll need to read, calibrate and evolve it on their own; or we help an already-structured company move toward an enterprise platform, but only once the organizational conditions to sustain it actually exist. The choice doesn’t start from the technology: it starts from whoever will have to live with it.

The stated goal is as banal as it is ambitious: raise the company’s technological level by one notch, and make sure that notch holds even when nobody’s watching. Which, as far as we’re concerned, is the only working definition of sleeping soundly at night.


Sources: Clusit 2026 Report; Verizon Data Breach Investigations Report (SMB figure: 2025 edition); 2025 SANS Detection and Response Survey; Microsoft / Omdia, State of the SOC 2026; ISC2 Cybersecurity Workforce Study 2025; Cybersecurity Insiders, 2025 SOC Survey.

← All articles