How to Choose a Cyber Range for Your Security Team

CT
CyberDefenders Team
Share this post:
CyberDefenders blog cover: "How to Choose a Cyber Range for Your Security Team" with an evaluation checklist.

To choose a cyber range, define what your team must be able to do first, then judge vendors on six things: realism of evidence, hands-on execution with real tools, team capability, threat relevance, scenario quality, and proof of analyst performance. Ignore scenario counts. Confirm everything in a pilot with your own analysts and your own tools before you sign.

Every vendor demo you will sit through this quarter shows the same thing: a ransomware scenario, a heatmap, a dashboard with green bars. After the third one they blur together, and the numbers start doing the deciding. Eight hundred scenarios versus three hundred. Fourteen categories versus nine. That is how teams end up buying a content library with a login page, and finding out twelve months later that nobody uses it and the dashboard cannot answer the CISO's one question: are we better than we were?

This is a guide to choosing without falling into that trap. It starts with your requirements, because a range is only good relative to what you need it to do. Then: the six criteria that separate platforms in practice, the questions that make vendors show rather than tell, a pilot design that can actually fail, and a scorecard you can drop into an RFP. It does not rank vendors; for that, see the best cyber range platforms compared. For the category basics, see what a cyber range is.

Start with your requirements, not the feature list

The first step in choosing a cyber range is writing down, in one page, what the team must be able to do after a year of using it that it cannot do today. Every later judgment about realism, features, or price is made against that page, and without it the vendor's feature list becomes your requirements by default.

Training objectives

Be specific about the outcome, not the activity. "Analysts complete 40 labs a year" is an activity and any platform can deliver it. "Every Tier 1 analyst can triage and correctly escalate a credential-theft alert within 30 minutes by month six" is an outcome, and only some platforms can get you there and show you they did. Write three to five objectives at that level of specificity. Typical ones: shorten onboarding to shift-readiness, close a named domain gap (cloud forensics is the common one), rehearse the two or three incident types leadership fears most, produce a readiness report leadership will accept. If you need to argue the objective before you can argue the budget, the hidden cost of an under-skilled security team is the framing that tends to land.

Individual vs. team capability

Decide how much of the problem is individual skill and how much is coordination. A team of strong analysts who have never worked an incident together has a coordination problem, and a platform that only sells solo labs will not touch it. A team of new hires has an individual skill problem, and paying for live-fire team exercises before the fundamentals are in place wastes them. Most teams need both, in a ratio that shifts over time; know your ratio before you look at demos.

Which functions have to be covered

List the roles the platform must serve: SOC tiers, incident response, DFIR, threat hunting, detection engineering, cloud security. Then be honest about depth. Many platforms cover every function at the entry level and one or two at depth. If your IR team is the reason for the purchase, the IR content is what you evaluate, not the catalog total. For the function-specific requirements behind this list, see our guides to a cyber range for SOC teams, a cyber range for incident response training, and a cyber range for threat hunting training.

The six criteria that separate cyber range platforms

Cyber range platforms differ on six things that matter and dozens that do not. Realism, hands-on execution, team capability, threat relevance, scenario quality, and evidence of performance decide whether the training transfers to real incidents. Scenario count, category count, and interface polish do not.

Criterion The question it answers The evidence that settles it
Realism Does the evidence look like production evidence? A senior analyst opens a scenario and reads the raw evidence before the questions
Hands-on execution Do analysts work in the tools they use on shift? Proportion of scenarios in production-class tooling, named tool by tool
Team capability Can several analysts work one incident together? A live multi-analyst scenario in the pilot, with a handoff and a debrief
Threat relevance Is the content current and aimed at your adversary? Release history with dates; ATT&CK mapping as an exportable file
Scenario quality Is a single exercise well built? Three scenarios at three difficulty levels, worked in full
Evidence of performance Can you see who can do what, and whether it is improving? An anonymized 90-day report from a customer team of your size

Realism

Realism means the evidence the analyst works is what production evidence looks like: real or high-fidelity packet captures, memory and disk images, event logs, cloud audit trails, with the noise, gaps, and dead ends that real environments produce. Synthetic, tidy evidence trains analysts to expect tidy evidence.

Hands-on execution

Hands-on execution means the analyst does the investigation with the tools they use on shift: a SIEM, EDR telemetry, forensic parsers, query languages. Not a simplified interface that mimics them, and not a walkthrough with a "next" button. Tool familiarity is part of what transfers.

Team capability

Team capability means several analysts can work one incident together, with handoffs, escalation, and a debrief. Not several analysts working the same solo lab in parallel, which is what most platforms mean when they say "team." Coordination is usually the first thing to fail in a real incident, and it can only be trained where several people share one investigation, so a platform that cannot host that is not addressing the problem at all.

Threat relevance

Threat relevance means the scenarios reflect the adversaries and techniques your sector actually faces, expressed in a framework you already use so coverage can be planned. In practice that is MITRE ATT&CK mapping you can export, plus a release cadence that tracks current threat reporting.

Scenario quality

Scenario quality is the craft inside a single exercise: a brief that reads like a ticket, questions that require finding evidence rather than guessing, a ground truth that is defensible, and a debrief that explains not just the answer but the path. It is the thing scenario counts hide, because craft does not aggregate.

Evidence of performance

Evidence of performance means the platform tells you what each analyst found, missed, and decided, and rolls that up per team over time, in a form a manager reads weekly and a security leader reads quarterly. Completion percentages are not evidence of performance. Neither is a leaderboard. For turning that reporting into a business case, see measuring cyber range ROI.

Evaluating the platform behind the training

Once the training passes, the platform has to pass procurement, security review, and a year of daily administration. These criteria rarely decide a purchase and frequently sink a deployment.

Integrations

Confirm which integrations are live today, not on the roadmap: single sign-on, automated provisioning and deprovisioning, an LMS or HR connection if training records have to land somewhere, and an API if you plan to pull data into your own reporting. Ask to see each one working in the trial tenant. "Supported" and "shipping" are different words.

If you want a standard vocabulary for the roles rather than your own job titles, the NICE Workforce Framework for Cybersecurity names work roles including Defensive Cybersecurity, Incident Response, Digital Forensics, and Threat Analysis, which travels better through HR and procurement than internal tier names do.

Deployment and security

A cyber range uses its own scenario data, so nothing sensitive should need to leave your environment, but the platform will hold your team roster and performance data, and that is enough to warrant your standard third-party review. Ask for security attestations (SOC 2 Type II, ISO 27001) as documents, not logos; in our experience few vendors publish them on their websites, so the request is normal and the answer is informative. Confirm data residency, retention, and what the vendor sees about your analysts' activity. On deployment model, cloud-hosted ranges deploy in days and scale on demand; on-premises makes sense where an air gap is mandated, or where you are validating your own infrastructure rather than training people. For capability-by-capability detail, see what to look for in an enterprise cyber range platform.

Administration

Ask who administers the platform day to day and how many hours a week it takes. You want team and role structures that match your organization, the ability to assign scenario paths to a cohort, and manager views scoped to what each manager owns. If the answer to "how do I see my Tier 2 analysts' progress this month" involves an export, the platform will not be used.

Support

Ask what support looks like after the sale: a named contact or a queue, response times, whether there is help for designing the program rather than only fixing the login. Ask the reference customer, not the salesperson.

Content updates

Ask for the release history for the last twelve months as a list with dates, and the roadmap for the next six. New labs weekly is a stronger signal than a large back catalog, because it tells you whether the library will still be relevant in year two of a three-year contract.

8 questions to ask cyber range vendors and providers

Ask these in writing and keep the answers. Vendors that show rather than tell are the ones to keep talking to.

  1. Open one scenario and show me the raw evidence before the questions. How was it produced?
  2. Which tools do analysts work in, and how much of the library runs in them?
  3. Which threat reports drove your most recent scenarios, and when did they ship?
  4. Can I get your MITRE ATT&CK mapping as a file?
  5. Show me an anonymized team report after 90 days. Can it tell me who can run an investigation today and where we are weak?
  6. How often do new labs ship, and what shipped in the last twelve months?
  7. Which integrations are live today, not on the roadmap?
  8. How is pricing structured, and what happens when headcount changes?

Run a pilot that can fail

A pilot is worth running only if it is designed to produce a no. Decide the success criteria, the participants, and the decision date before it starts, and use your own tools and analysts, not the vendor's demo team.

A working design takes four to six weeks. Choose five to ten analysts across tiers, including at least one skeptic and one senior who will judge realism. In week one, every participant works the same three baseline scenarios; that gives you a first read on evidence quality and on the report. Weeks two to five run the weekly cadence you intend to keep, one scenario per analyst matched to their tier, plus one team exercise if team capability is a requirement. In the final week, repeat the baseline and hold a review. Before the pilot starts, write down what passing looks like. Five criteria are usually enough:

  • The senior analyst rates evidence realism as acceptable.
  • Managers can read the reporting without being trained on it.
  • SSO and provisioning work in the trial tenant.
  • The team exercise produced a debrief with at least one coordination finding.
  • The price model has been confirmed in writing.

Any criterion missed is a no, or a renegotiation. Deciding that after the pilot is how teams talk themselves into a purchase.

We run pilots to this design for teams evaluating our own platform, and we would rather lose one than have a team buy on a demo. A pilot that cannot fail is a procurement formality, and everyone in the room knows it.

RFP checklist and scorecard

Put the criteria above into a weighted scorecard so the evaluation team scores the same things from the same evidence. The weights below reflect what predicts whether the platform is still in use in year two; adjust them to your requirements page.

Criterion Suggested weight Evidence to request Score (1 to 5)
Realism of evidence 20 Senior analyst works three scenarios; vendor states evidence provenance  
Hands-on execution 15 Proportion of scenarios in production-class tools; analyst feedback from pilot  
Team capability 15 Live multi-analyst scenario in pilot; handoff and debrief observed  
Evidence of performance 15 Anonymized 90-day customer report answers the three questions  
Scenario quality 10 Three scenarios at three difficulty levels, worked in full by your analysts  
Threat relevance and content cadence 10 Release history with dates; exportable ATT&CK mapping  
Function coverage (SOC, IR, DFIR, hunting) 10 Depth in the functions on your requirements page, not catalog total  
Enterprise fit (integrations, security, administration, scalability) 5 Live integrations in trial tenant; attestation documents; admin hours estimate  

Score each vendor independently, then meet to reconcile differences of two points or more; those gaps are where the real evaluation happens. A vendor with the highest total but a 1 on realism or assessment should still lose, because those two criteria are the ones the others cannot compensate for.

Start with a pilot, not a contract

Whatever you do with the rest of this, do not sign on a demo. The scorecard above is worth more to you filled in from a month of your own analysts' work than from four vendor presentations, and the vendors worth buying from will prefer it that way.

If your requirements page points to blue-team investigation skill, per-analyst evidence, and a team that has to train together, talk to the CyberDefenders enterprise team and ask for the pilot described above rather than a presentation. Score us on the same card as everyone else.

FAQ

How long should a cyber range evaluation take?

Six to ten weeks from requirements to decision: one to two weeks to write the requirements page and shortlist, one to two weeks of demos and written questions, four to six weeks of pilot, and one review meeting. Shorter than that usually means the pilot was skipped, and the pilot is the only part that produces evidence.

Should we choose a cyber range by number of scenarios?

No. Scenario count says nothing about realism, tool fidelity, or whether the library is current. Judge three scenarios at three difficulty levels in depth, ask for the release history with dates, and weight your scorecard toward realism and assessment. A hundred well-built, current scenarios beat a thousand thin ones.

What is the biggest mistake teams make when choosing a cyber range?

Buying on the demo and skipping the pilot. Demos are run by the vendor's best presenter on the vendor's best scenario. A pilot puts your analysts, your tools, and your skeptic in front of the platform for a month with written pass criteria, and it is the only step that can tell you no.

Can a small security team use this process?

Yes, in a lighter form. A team of five still needs the requirements page and still needs to judge realism and assessment, but the pilot can be two weeks with three analysts, and the integrations and administration criteria matter less. The five core criteria do not change with team size; only the weight on enterprise fit does.

Tags:soc trainingsecurity analyst trainingCybersecurityBlue Team