10 Questions to Ask When Choosing a Managed Service Provider
Downtime costs $4,537 a minute while the SLA penalty is largely symbolic. Ten questions that show what lies behind the price, and how to weigh the answers.
Short answer: The price line does not show the difference between two proposals. These do: what the SLA measures, who answers the phone at 3 a.m., what happens when it is breached, and how you get your data back when you want to leave.
In putting this list together we tried not to leave out the questions that make us uncomfortable too. Publishing them as a provider means being asked the same things, which is the point.
1. What exactly does the SLA measure?
"99.9% uptime" on its own means nothing. Ask what counts as working: is the server up, is the service responding, or is the application doing its job? A server can answer a ping while the application has crashed, and under a narrowly defined SLA that is not a breach.
Ask whether planned maintenance windows count against it. In most contracts they do not.
2. Is the commitment to first response or to resolution?
The gap between the two is large. "15-minute response" can mean a ticket was opened. So ask whether there is a commitment on time to resolution for a critical incident, or whether only the acknowledgment is measured.
3. What happens if the SLA is breached?
Is there a penalty clause, and how does it work? Usually it takes the form of a credit on the next invoice, and the amount is largely symbolic. That is not necessarily bad, but sign knowing it will not cover the cost of a real outage.
For scale: PagerDuty's 2026 data puts the average cost of downtime at $4,537 per minute, so a three-hour incident exceeds $800,000 on average. Ten percent of a monthly service fee is a rounding error next to that. The penalty's real value is as a signal of how much the provider trusts its own commitment.
The more important question: when a breach occurs, do you have to notice, or does the provider report it? Providers that report their own breaches are rare, and they are the ones you want.
4. Who is watching at 3 a.m.?
Ask about the on-call model. Does a call center answer, or does the call go straight to an engineer? Does the person on call know your systems, or are they reading a procedure?
The most revealing question: is the team that responds to incidents the same team that built the system? If the answer is no, you will pay for a learning curve on every incident.
Add one more: can they decide on a rollback or a failover without your approval, and if not, who on your side holds that authority? In the incident escalation matrix, we calculated what the authority gap costs in minutes.
5. How do I get access to logs and records?
When you go into an audit, or investigate an incident, the records will be in the provider's system. Ask: how long do they retain them, how many hours do they need to produce them on request, and is their own team's access to your systems logged too?
If it is not in the contract, the gap is yours at audit time. We covered how long each record must be kept in log retention periods, and the banking requirement on where it must be kept in BDDK treats cloud as outsourcing.
6. What happens if I want to leave?
Asking this during a sales conversation feels awkward, which is exactly why it should be asked. In what format, and how quickly, do you get your data back? Do configuration and documentation transfer to you? If you depend on tooling the provider built in-house, does the environment run without it?
A contract that is easy to exit also gives the provider more reason to deliver well.
7. What is in scope, and what gets billed separately?
The out-of-scope list should be in writing, just like the in-scope list. The items that most often cause disputes: new server builds, version upgrades, project work, recovery effort after a major incident, and advisory requests.
8. How are backup and recovery proven?
Do not ask whether backups run. Ask whether restores work. Are periodic restore tests performed, are the results reported to you, and when was the last one?
Without a report, backup is an assumption. Veeam's 2025 research shows how common that is: only 10% of ransomware victims recovered more than 90% of their data, and average recovery took 24.6 days. Almost all of them had backups.
Ask for two concrete things: the date of the last restore test and the recovery time it measured. We explain what makes a test meaningful in the disaster recovery failover test guide.
9. Who are the subcontractors?
Your provider may have someone else delivering part of the service: the data center, the network, a specific area of expertise. Ask for a list of subcontractors and to be notified when it changes. You do not want to discover where the chain of responsibility breaks only after an incident.
10. Who will look after me, and for how long?
Is there an assigned technical lead? How is handover done when the team changes? If turnover at the provider is high, nobody who knows your environment will remain, and you will find that out during a crisis.
How to weigh the answers
Finding a provider with a clean answer to all ten is hard. The answers do not have to be perfect, but they must not be vague. "We'll see", "generally" and "it depends" mark the items that will turn into disputes after signing.
One more thing: ask for the answers to go into the technical annex of the contract. What was said in a meeting carries no weight in an audit or a dispute.
Sources
- PagerDuty 2026 MTTR data: cost per minute of downtimeVeeam 2025 Ransomware Trends: recovery rates and times actually achievedHow we do this work: managed services
This list was compiled from the questions we encounter in our own customer conversations as a managed service provider. We suggest asking Eclit the same questions.