Torq advances defensive AI through research in security simulation, model post-training, agentic harness optimization, and customer specific learning. We combine proprietary models and research infrastructure with leading frontier models to build the systems behind autonomous security operations.

A model doesn’t defend anything. A system does.
The hard problem in security AI isn’t a benchmark score alone. It’s producing trustworthy decisions inside a live SOC — on your data, against real adversaries, every hour of the day. That takes more than a model. It take a harness that wields it, data that grounds it, and a research program that develops and evaluates them together.
our approach
Purpose-Built Intelligence, Engineered End to End
Torq invests across the models, data, simulation environments, retrieval systems, and agentic harnesses that autonomous defense depends on. These are distinct research directions, connected by proprietary infrastructure, curated security knowledge, and systematic evaluation, rather than reliance on any single model.
The Right Model for Each Task
A hybrid of proprietary Torq models and leading frontier models, with routing optimized for the requirements of each investigation step.
Optimized with the Harness
Models and agentic harnesses are tuned together, not in isolation — because in production, the harness is what turns a model into a defender.
Forged in a Digital Twin
Torq’s Cyber Range supports model training, harness development, and evaluation against realistic enterprise and MSSP security scenarios.
Adapted to Every Environment
Every environment shapes the AI: Reflex models, Recall of historical cases, organizational context, guidance, and agentic memory all reflect how each SOC works.
the cyber range · security simulation
A Digital Twin of the SOC We Defend
Torq’s Cyber Range is a high-fidelity digital twin of enterprise and MSSP security operations, built and operated by experienced practitioners. It brings multiple threat sources into a controlled environment for developing and evaluating defensive AI against the conditions of a working SOC.
- TWIN – Enterprise and MSSP fidelity: Identity, endpoint, cloud, network, and SIEM layers modeled as they appear in production.
- ADVERSARY – Real attack scenarios, many sources: Live TTPs and campaigns replayed against the twin, continuously refreshed.
- HUMAN – Curated by practitioners: Security operations veterans define, label, and validate scenarios.
Product outcome: Models and agents can be exercised against representative SOC environments before deployment, helping identify investigation gaps and failure modes before they reach a live security operation.
Multiple adversary sources are replayed into a representative SOC twin; practitioners curate and validate scenarios.
Normal activity and synthetic attacks merge into telemetry with scenario-derived ground truth.
Torq Forge · Synthetic Telemetry
Realistic Security Telemetry, Generated at Scale
Real-world attack data can be sensitive and difficult to label. Torq Forge generates realistic user, network, and device activity at scale: normal business operations interwoven with synthetic attack patterns modeled on real adversary behavior. It supplies the Cyber Range with controllable, reproducible telemetry and scenario-derived ground truth for defensive AI research.
- NORMAL – Lifelike baseline activity: Users, hosts, and services behaving as they do on an ordinary day, the signal beneath the noise.
- SYNTHETIC – Attacks based on real TTPs: Adversary patterns injected into the stream with labels derived from the underlying attack scenario.
- SCALE – Controllable and reproducible: Configurable scenarios and telemetry volumes, with repeatable conditions for controlled research comparisons.
Product outcome: Torq can train and test defensive AI against attack patterns beyond those recorded in an individual customer’s environment.
Tagged Datasets · Research Data
Curated Fuel for Three Engines
Inside the Cyber Range, Torq’s security research team maintains tagged datasets as labeled records of activity, alerts, decisions, and outcomes.
- Post-training: Specializing models on security reasoning, tools, and the language of real cases.
- Harness optimization: Feeding the loop that evolves prompts, retrieval, and step-by-step model routing.
- Evaluation frameworks: Held-out, versioned benchmarks that decide whether a change actually ships.
Product outcome: Research changes are assessed against expert-labeled benchmarks to measure gains and identify regressions before release.
A curated data foundation supports post-training, harness optimization, and held-out evaluation.
Research direction: learned investigation policies, with objectives spanning consistency, investigation quality, and inference efficiency.
Post-Trained Models · Active Research
Orchestration Models, Built to Investigate
The hardest part of an investigation isn’t any single lookup — it’s the orchestration: what to check, in what order, when to pivot, when to escalate, and when the evidence supports closing a case. Torq’s post-training research targets this investigation policy directly, developing purpose-built models designed to learn through repeated investigations in the Cyber Range.
- RL GYM – The Cyber Range as an RL gym: The Cyber Range provides the foundation for reinforcement-learning research into end-to-end investigation policies, with objectives centered on verdict quality and efficient evidence gathering.
- DENSE – Multimodal, dense architectures: Our research explores multimodal, dense models for orchestration, grounded in logs, alerts, entities, and investigation timelines.
- FOCUSED – Specialization beyond prompting: Post-training targets investigation behavior in model parameters, not just instructions in a prompt. The objectives are more consistent orchestration, stronger investigation performance, and lower inference cost.
Research objective: More consistent investigations and lower inference cost through orchestration models purpose-trained for defensive security.
Loop Engineering · Harness Optimization
An Automated Research Loop for the Harness
Torq runs an automated research loop that proposes, tests, and selects improvements to the agentic harness. Drawing on updated tagged datasets, it generates candidate changes using models across authorized model gardens, compares them on security-specific tasks, and advances selected improvements through evaluation-gated releases.
- Prompt evolution: Algorithmic search over prompt variants — keeping the phrasings that raise measured performance.
- Data reduction: Transforming retrieved data to preserve relevant security context while reducing the tokens needed to represent it.
- Step and model routing: Evaluating how investigation steps are decomposed and routed across models to balance task performance and cost.
Product outcome: Prompt, retrieval, and routing improvements are developed systematically and assessed against security-specific evaluations before release.
Generate → optimize → evaluate → release.
Research candidates advance through evaluation gates.
Recall supplies typed graph context and case history for evidence-grounded investigation.
TORq recall · structured retrieval
Structured Retrieval for Security Reasoning
Torq Recall is a structured retrieval layer built for security operations. It returns typed context grounded in the Torq Context Graph, bringing entities, relationships, and case history into the agentic harness in a form designed for security reasoning.
- TYPED – Structured over scattered: Entities, relationships, and case history returned as structure the harness can act on.
- GROUNDED – Grounded in the Context Graph: Environmental context and historical cases provide evidence for the current investigation.
- LEAN – Signal per token: Co-designed with Loop Engineering’s data-reduction track to keep context dense and relevant.
Product outcome: Investigations use structured environmental context and case history, providing a clearer basis for reviewing the evidence behind a verdict.
torq reflex · customer-specific learning
A Model that Learns Your SOC — Only Yours
Two SOCs are never alike. Beyond shared models, Torq trains a dedicated encoder model with a classifier head for each customer environment, learning from that customer’s alerts and cases to reflect its operational practices, priorities, and verdicts. This adapts a model through training, rather than retrieval alone: Torq Reflex learns customer-specific decision patterns in its parameters.
- PER-ENV – Trained on your data, for you: Each environment gets its own encoder + classifier head. Fully isolated, never pooled across customers.
- LEARNS – Your definition of what matters: The model absorbs how your team actually triages, escalates, and closes cases.
- IMPROVES – Feedback-driven model updates: Analyst-confirmed verdicts and corrections provide additional training signal for subsequent model updates.
Published comparison: 92% vs 78% agreement on analyst-correct alerts.
Product outcome: Your SOC gains a dedicated model trained on your team’s decisions, with analyst feedback informing subsequent updates rather than only being appended to a prompt.
Each customer gets an isolated encoder + classifier head, trained and updated on their own cases.
evaluation & validation
Measured, Not Asserted
Architecture and ambition need evidence. Torq evaluates AI capabilities against security-specific benchmarks drawn from production scenarios, labeled by experts, and checked against ground truth. Where trade-offs exist, malicious recall is a priority, evaluated alongside noise reduction, grounding, and execution quality.
| AI Capability | What We Measure | How It’s Validated |
|---|---|---|
| Alert Triage | Malicious recall, noise reduction | Expert-labeled production alerts across multiple EDR vendors, checked against ground truth. |
| Socrates — ActionPlan | Instruction adherence, factual grounding | Deterministic, step-by-step plan-execution checks over diverse playbooks. |
| Socrates — Conversational | Tool-execution accuracy, factual grounding | Security-specific tool-invocation benchmarks measured against expected outcomes. |
| Data Transformation | One-shot success rate | Anonymized production telemetry from real-world usage. |
| Torq Reflex | Corrected-verdict agreement | Comparison with a semantic-similarity baseline; see the published result below. |
Published result · torq reflex
Customer-Specific Learning, Measured
Torq Reflex matched corrected analyst verdicts on 92% of evaluated alerts, compared with 78% for a semantic-similarity baseline.
14 percentage points higher agreement in the reported comparison.
92%
Torq Reflex
78%
Semantic-similarity baseline
Torq-reported evaluation of analyst-corrected alerts, published June 25, 2026. This is not overall detection accuracy.
How We Evaluate
curated benchmarks
Real Scenarios, Expert-Labeled
Datasets drawn from real, production-grade security scenarios and labeled by experts, curated for quality over quantity.
ground truth
Validated, Not Guessed
Deterministic verification wherever an answer is objective; expert human review wherever judgment is required.
production telemetry
Confirmed in the Field
Aggregate, anonymized success metrics from real-world usage round out every benchmark.
Every benchmark is re-evaluated at least monthly — the most critical, weekly — with continuous production monitoring and fresh runs after any model or architecture change. The datasets keep growing as products evolve and the threat landscape shifts.
Built for enterprise trust
Ambition that an Enterprise Can Actually Adopt
Purpose-built AI only matters if a CISO can put it into production. Torq’s research program is engineered around the controls enterprises and MSSPs require.
Pre-Customer Isolation
Reflex models are trained and served per environment. Your data trains your model, never a shared one.
Authorized Subprocessors
Loop Engineering runs across model gardens — AWS Bedrock and Google Vertex AI — operated as authorized Torq subprocessors under enterprise data governance.
Evaluation-Gated Releases
Model and harness changes are assessed against held-out, versioned evaluations before release.
Auditability and Human Oversight
Agentic decisions are grounded in context and available for review, with human oversight of autonomous operations.
proprietary & frontier
We don’t choose between our models and the world’s best. We use both.
Torq combines leading frontier models with our own purpose-built models, optimizing their use across investigation tasks. Our proprietary research spans the simulation environments, curated datasets, customer-specific models, and agentic harnesses that turn model capabilities into defensive security systems.
Frontier model partners:

OpenAI
Anthropic
Model gardens and infrastructure: AWS Bedrock Google Vertex AI Authorized Torq Subprocessors
The Future of Security Operations is Agentic
See it in action.

