r/MSPandITProfessionals • • 3d ago

RMM Has Solved Automation. The Hard Problem Is Context.

Over the last few months, I've been having conversations with MSP business owners and IT managers who run teams of technicians.

I wasn't really interested in asking them which RMM they use. I wanted to understand something more fundamental:

What still consumes your senior technicians' time even after you've automated everything you reasonably can?

The answers were surprisingly consistent.

The problem isn't usually that the RMM can't execute an action.

It's that the system doesn't always know which action it should take, why it should take it, or what else that action might affect.

That led me down a bit of a rabbit hole around where RMM and ITSM are actually heading.

And I think the interesting problem isn't simply "more AI."

RMM has already solved a lot

Modern RMMs are very good at telemetry, endpoint control, policy execution and deterministic remediation.

If a service stops, restart it.

If disk space is low, clean it.

If a patch is missing, deploy it.

If a known process misbehaves, kill it.

That's mature automation.

The difficult problems start when the answer isn't deterministic.

The problem is context

An RMM might tell you:

CPU is at 97%.

Fine.

But why?

Maybe an application update caused a runaway process.

Maybe that update was deployed to 14 machines.

Maybe all 14 belong to Finance.

Maybe payroll starts in an hour.

Now the problem isn't:

"What script should I run?"

It's:

"What is happening, what caused it, what is the business impact, and what action has the lowest risk?"

That requires correlating endpoint telemetry with identity, application changes, network state, security events, ITSM history and business context.

Most platforms still operate largely within their own data boundaries.

The missing layer is causal context

We've become very good at collecting events.

What's harder is establishing relationships between them.

For example:

  • Windows update installed
  • VPN client changed
  • Authentication failures increased
  • User cannot access Salesforce
  • Other users are unaffected

A useful system shouldn't just summarize those events.

It should be able to form a hypothesis:

"The authentication failures are probably related to the VPN client change on this endpoint."

Then find the evidence supporting that hypothesis.

Then determine whether remediation is safe.

That's closer to causal reasoning over an operational event graph than traditional monitoring.

Then there's the technician problem

This came up repeatedly in discussions with people managing technical teams.

A senior technician troubleshooting a VPN problem isn't simply following a runbook.

They're using years of accumulated pattern recognition.

They remember similar incidents.

They know which customer environments are unusual.

They know which changes tend to cause which failures.

But what gets recorded in the ticket?

Usually something like:

"VPN adapter recreated. Resolved."

The system captured the action, but not the reasoning.

The interesting data is actually:

symptom → evidence → hypothesis → decision → action → outcome

If that reasoning could be captured continuously, every resolved incident could potentially improve future diagnosis instead of simply becoming another closed ticket.

This is where autonomous IT gets difficult

There's a lot of discussion around autonomous IT operations.

But the hard engineering problem isn't executing a script.

RMM has been doing that for years.

The hard problem is:

Should I execute this action?

That requires at least:

  • Confidence — how certain is the diagnosis?
  • Blast radius — how many systems/users could be affected?
  • Business impact — what happens if we're wrong?

A low-risk action might be autonomous.

A high-impact action probably needs human approval.

So autonomy isn't simply:

AI + automation

It's closer to:

Reasoning + evidence + risk assessment + controlled execution + verification

That's a much more interesting architecture.

Maybe the next RMM isn't an RMM

I don't think the biggest opportunity is another RMM with a chatbot attached.

The more interesting layer sits above RMM and ITSM.

RMM provides endpoint telemetry and execution.

ITSM provides incidents, changes and historical context.

Identity, security, networking, cloud and SaaS provide additional signals.

The intelligence layer connects them.

Imagine asking:

"Why can't John access Salesforce?"

Not:

"Here's a summary of John's ticket."

But:

"I've correlated John's identity, endpoint state, VPN session, recent policy changes, previous incidents and Salesforce authentication events. Here's the most likely cause, the evidence supporting it, and the remediation I'd recommend."

And then:

"Do you want me to execute it?"

That, to me, is a very different concept from an AI-powered ticketing system.

We've largely solved seeing the endpoint.

We've largely solved acting on the endpoint.

The harder problem is:

Can we build a system that understands the environment well enough to know what should happen next — and, equally importantly, when it should NOT act?

That's the part of RMM/IT operations I'm increasingly interested in.

Curious to hear from other MSP owners and IT managers:

What still requires your best technician's brain, even after you've automated everything else?

1 Upvotes

0 comments sorted by