VORTRAQ
Follow on X ↗
Trading research · validation · execution

Build the system you actually want to test.

Start with a signal and trade it. Add filters if you want them. Tear the population apart. Ask whether the venue changes it. Train ML if it helps. Vortraq gives you the depth without making the depth compulsory.

Vortraq supports trading across 1m, 5m, 15m, 1h and 4h timeframes, with cross-timeframe information available to research and machine learning.

Built independently in the UK. Vortraq was created by Tadas Siksnius and is developed and operated by Vortraq Ltd. The product is presented on its software, methodology and reproducible evidence rather than founder credentials.
The actual application. No concept render.
Product map

Open the part you care about.

Vortraq is too large for one endless homepage. Treat the site more like the application itself.

STRUCTURE

Viewer & Signals

Range Bands and Parallel Channels are structural engines, not decorations. See the geometry and state that existed when the candidate fired.

OPEN →
POPULATION

Trade Policy

Replay filters, ARM1 / ARM2 admission profiles and your own rules. See the exact population each decision leaves behind.

OPEN →
FORENSICS

Trade Analysis

Dozens of analytical families can interrogate the same recorded population, then push attractive findings into chronological tests instead of trusting the first green number.

OPEN →
VENUE BEHAVIOUR

Hyperliquid Lab

Replay the same population as one net position and see whether the venue helps, hurts or changes the strategy.

OPEN →
LEARNING PROBLEM

ML Workshop

Search the learning problem, not just the algorithm: population, target, feature families, history, retraining, model head, operating point and OOS behaviour.

OPEN →
MODEL BEHAVIOUR

ML Lab

Direction and MFE answer different questions. Test whether either helps alone, whether the second adds anything, and what their composition actually does to the trades.

OPEN →
EXECUTION

Live Trading

Connect an account, choose signal and timeframe, trade. Maker Hybrid is an optional execution mode. It is not required to use Vortraq.

OPEN →
VALIDATION

Standard tests first. Parity goes further.

Walk-forward, holdout/OOS, CPCV with purge and embargo, time-shift, permutation and Monte Carlo are part of the validation toolkit. Live/offline parity asks an additional question: did research reconstruct what live actually saw?

OPEN →
THE PART THAT TOOK THE WORK

The layers are not separate toys.

The same identifiable decision can start in a structural engine, be admitted or rejected by Trade Policy, reappear in Trade Analysis, become learning material in Workshop, be scored by Direction or MFE, be replayed under Hyperliquid's one-position behaviour and later be checked against what the live engine actually saw.

STRUCTURE
POPULATION
ANALYSIS / ML
VENUE
LIVE / PARITY

You do not have to export a backtest from one tool and hope the next tool rebuilt the same trade.

Viewer & Signals

Keep the numbers attached to the market.

A candle is only the beginning.

Vortraq can carry Range Band geometry, Parallel Channels, regime, RSI, volatility, structural timing, event history and policy state with the candidate.

Parallel Channels are a full structural engine in their own right. Their geometry, formation, position, touches and breaks can stay attached to the decision and travel into later analysis or learning instead of becoming a line somebody drew on a chart.

The point isn't to use everything. It's to be able to ask whether any of it matters.

Viewer ties the recorded decision back to the structure that produced it.
Trade Policy

Change the population before you change the model.

Filtering can win twice.

A filter may improve the trades directly. Or it may remove a noisy part of the population and make the remaining ML problem easier to learn.

ARM1 and ARM2 are built-in structural admission profiles for exactly this kind of experiment. Use them, ignore them, or build your own rule set — then replay the before/after population rather than guessing what the filter changed.

Filtering isn't only about finding better trades.

It can be about creating better learning material.

Replay the rule, compare before and after, export the resulting population if it deserves another experiment.
Trade Analysis

Finding something good is the easy bit. Believing it is harder.

Trade Analysis can uncover relationships inside a fixed, already-recorded trade population. The first attractive number is not treated as the answer. Change the bucket definitions. Split the period. Move through regimes. Push the relationship into chronological testing. Look at the interactions around it. If it disappears as soon as you stop asking the question in exactly the right way, that is useful evidence too.

70+analytical modules
154,903example report lines
13,664trades resolved
1M+state rows loaded

Research map

Click a branch. This public view shows the shape of the engine without publishing the research recipe.


Trade Analysis

Some sections discover. Others deliberately try to overturn what looked convincing — across time, regimes, nearby definitions and competing explanations.

TRADE
ANALYSIS

Decision-time studies keep future movement, MFE and outcome on the label side of the experiment.

Chronological FIT → TEST guardrails can reject a cohort rather than manufacture a conclusion.

Before report generation

Your buckets don't have to be our buckets.

“High”, “low”, “early”, “strong” or “stretched” do not have to become sacred because one definition happened to produce a pretty result.

Change the selected floors and bucket definitions, then run the same population again. If a relationship only survives one convenient boundary, you have learned something useful — just probably not what you hoped.

Vortraq gives you more ways to look for fragility before you trust a result. It cannot stop you ignoring what you find.

Hyperliquid Lab

The venue can change the strategy.

Hyperliquid does not keep independent long and short trades on the same market. They interact inside one net position. Hyperliquid Lab takes the population you already have and asks what that changes.

RAW OR TRADE POLICY POPULATION
INDEPENDENT BOOK
/
HYPERLIQUID NET POSITION
COMPARE WHAT CHANGED
Vortraq Hyperliquid Lab comparing independent-trade and one-net-position execution methods
One accepted signal population, replayed through multiple execution interpretations. The Lab compares the independent reference with Hyperliquid Native and research modes without rewriting the source population.

Same population. Different economic sequence.

Same-side signals can add to the existing position and change its average entry. Opposing signals can reduce, flatten or reverse exposure instead of opening a separate trade.

Sometimes the TP/SL is no longer the event that ends the trade.

That is why the Lab is useful both on a final candidate population and on an ugly pile of raw signals. It does not decide what belongs in the population. It shows what one-position execution does to it.

1net position
opposing collisions
natural completion
Δvenue effect
Venue fit check

Is Hyperliquid better or worse for this exact setup?

Compare the existing independent-trade reference with Hyperliquid's one-position interpretation. Collision rate, natural exits, realised PnL, fill count and fee drag tell you where the difference actually came from.

Trade Policy can shape the population first. Or leave it completely raw and let the Lab expose what the venue does without filtering.

HYPERLIQUID NATIVE

Do nothing clever first.

Replay the accepted signal stream under the venue's net-position mechanics. This is the baseline.

RESEARCH D

Opposing Exit

Ask whether an opposing signal is more useful as the exit of the position already open than as a separate new trade.

RESEARCH E1 / E2

Progression experiments

Test whether same-side structural progression should bank the position or carry the position forward under a changed exit schedule.

These are research modes, not Trade Policy rules and not ML labels. They do not rewrite the population that entered the Lab.

WHAT IT MODELS

Net-position semantics.

Recorded dataset prices are replayed through one-position accounting so stacking, opposing reduction, flattening, reversal, average entry and fill-count effects can be studied.

WHAT IT DOES NOT MODEL

Historical Hyperliquid market microstructure.

It is not an order-book reconstruction and does not claim historical fill probability, funding or slippage fidelity. Those belong to a different layer of evidence.

Hyperliquid Lab is a venue-behaviour model, not a promise that a historical order would have filled exactly that way.

ML Workshop

Choose the problem the model gets to see.

The learner is only one choice. Workshop lets the population, target, feature families, amount of history, retraining cadence, model head, operating point and chronological validation move independently so you can find out which part actually mattered.

Search is not proof. A configuration that wins a search is a candidate, not a conclusion. Vortraq keeps discovery and the harder question separate: does the relationship survive outside the conditions that selected it?

What the search does not do

It does not correct for how many candidates it evaluated.

Comparing many candidates and keeping the winner introduces selection risk. Vortraq does not pretend that OOS turns unlimited experimentation into proof, and it does not currently apply a best-of-N statistical correction to the winning search result.

The evaluation count belongs with the result because it changes how much confidence you should place in a winner. Search finds candidates. Forward evidence on data that did not exist when the search ran remains the stronger test.

Configuration

The tests are parameters, not verdicts.

Population, target, history length, retraining cadence, model head, operating point and validation regime can move independently. That freedom is deliberate, but not every combination answers the same question.

A holdout that freezes a model across a period far longer than your normal retraining cadence is testing stability under model ageing; it is not reproducing the same process as a frequently retrained walk-forward. Six years of history with weekly retraining can be a perfectly valid experiment, but it is answering a very different question from a long frozen holdout over the same span.

Embargo length must make sense for the trade horizon being tested, and a screening-resolution permutation run should not be mistaken for a final statistical verdict. Vortraq runs the research configuration you choose. It does not pretend every possible combination means the same thing merely because the run completed.


XGBoost · CatBoost · LightGBM · SGD

The model heads are useful. The experiment around them is the point.

POPULATION

What should the model learn from?

Start with all eligible trades, then decide which population belongs in this experiment. Filtering can improve the trades directly, but it can also change the quality of the learning material.

ALL TRADES
FILTER / POLICY
SELECTED POPULATION
ML

A completed search is a trading experiment.

Population size, feature selection, economic outcome, drawdown, consistency and retention all matter. A classifier score by itself is not the finish line.

Don't want to live in Workshop?

Once you're happy with the population and feature set, Vortraq can keep an accumulating training seed and retrain automatically as fresh trades arrive.

The current automatic path retrains after each additional block of 200 trades. That is an operating choice, not a claim that 200 is universally optimal.

ML Lab

Prove that the second model adds something.

Direction asks whether the trade eventually wins. MFE asks a different question about the excursion available after entry. ML Lab puts those learned views back onto the same population and asks the only useful question: does either one — or the combination — actually improve what you would keep?

A trained model isn't the finish line.

Utilisation asks what happens when you actually use the scores to retain or reject trades.

Direction and excursion are different questions. A second model only matters if it contributes information the first one did not already know.

Composition Grid

Put separate learned questions back onto the same population and see where they agree, disagree and whether the combination actually changes the outcome.

Policies

Reusable decision logic, when you want it.

CLOCK

Timing and age

Ask whether the candidate is still structurally timely.

RAINBOW V1 / V2

Structural permission

Weight surrounding evidence or require exact permitted setups.

STRETCH

Separate candidate path

Built around evolving band geometry rather than simply judging an existing signal. It can form its own candidate population and then pass through the same research, learning and validation machinery.

Policies are optional. They do not replace manual structural research, Trade Policy, Parallel Channels or ML.

Results & Simulation

A model metric isn't a trading result.

Translate the research into a trading path.

Look at costs, retention, drawdown and capital use rather than stopping at classifier metrics.

Would you actually want to live through this result?

Live Trading

Deep research. Very short route to live.

CONNECT ACCOUNT
CHOOSE SIGNAL
CHOOSE TF
TRADE

No ML model required. No prepared population required. No month in Trade Analysis required. Add those layers if you want them.

Execution is evidence too

A signal is not a fill.

Live execution is treated as its own stateful process rather than assuming a research entry magically became a trade. Orders, acknowledgements, fills, position state and protective exits belong to the live record; the journal records what actually happened rather than replacing it with the research assumption.

Signal source and execution venue are deliberately separate. Canonical market data keeps the decision engine consistent. Execution state comes from the connected venue. A Binance decision price is not treated as proof of a MEXC or Hyperliquid fill.

STATE / HEALTHY

Signal feed certified.

Canonical Binance candles current. New entries allowed.

STATE / INTERRUPTED

Don't pretend the missing data didn't happen.

Signal state becomes uncertified. New trade admission fails closed while the gap is dealt with.

STATE / RECOVERED

Restore the history. Rebuild the state.

Missing closed candles are returned through the same processing path before the signal state is re-certified.

MEXC

Opposing positions can coexist.

A later short does not have to erase the long already open.

HYPERLIQUID

Exposure is netted.

An opposing signal can reduce, neutralise or reverse the position already open. That can materially change realised performance even when the Vortraq signal stream is identical.

Hyperliquid Lab lets you measure that difference before live deployment. Live execution then supplies the venue evidence the simulator cannot manufacture.

Maker Only is optional

Vortraq does not require Maker Hybrid execution. It is an optional mode for traders who deliberately want maker-first entries and want the resulting fill behaviour to be part of the live experiment.

It is not the live architecture

With Maker Only disabled, Vortraq uses its normal supported execution path. Signal selection, Trade Policy and ML do not depend on Maker Only being enabled.

When Maker Hybrid is used, Vortraq does not turn a historical candle touch into an invented fill. Fill quality and venue behaviour remain execution questions, and the live record shows what the venue actually accepted and filled. Maker Hybrid is optional and is not required to use Vortraq.

Keys and custody

Your exchange account remains your exchange account.

Vortraq is execution software, not a custodian. It does not take custody of trading funds. Exchange connections should be created with only the permissions required for trading; withdrawal permission is not part of the trading workflow.

Venue credentials can be revoked from the exchange independently of your Vortraq research data. Exact credential-storage and connection details will be documented with the release build rather than hidden behind a generic “secure” claim.

Real audit output

What one Vortraq decision actually looks like.

A trade is not stored as a line saying SHORT → +0.8%. Vortraq records the decision-time market state, the resulting trade, the execution lifecycle and outcome measurements as they become available.

This is real Vortraq output, not a mock-up. The example below comes from an actual BTCUSDT 5m short. It is abridged only for display length; values have not been simplified into a marketing representation.

01 / candidate

The world as the decision saw it.

{
  "record_type": "candidate",
  "timestamp": "2026-08-25 16:39:00+00:00",
  "timeframe": "5m",
  "side": "short",
  "event": "rb_cross_dn",
  "bar_idx": 3606,
  "tf_idx": 721,
  "rb_pos": "high",
  "rb_zone": "upper",
  "band_order": "B>A>D>C",
  "rb_hier": "cross_dn",
  "rb_reg": "bull",
  "d_slope_bucket": "strong_down",
  "d_slope": -81.91769426602191,
  "A_lower": 79225.87394556717,
  "A_upper": 79343.6189115757,
  "B_lower": 79184.38601034538,
  "B_upper": 79388.3262738076,
  "C_lower": 78588.80005454503,
  "C_upper": 78877.21514104724,
  "D_lower": 78874.96544727162,
  "D_upper": 79247.30772287765,
  "atr_used": 168.7324102613563,
  "A_over_B": false,
  "B_over_C": true,
  "C_over_D": false,
  "pc_state": "none",
  "pc_event": "none",
  "q_score": 1.03,
  "q_pass": true,
  "rsi": 50.05958456607261,
  "rsi_zone": "neutral",
  "rsi_1m": 47.53570908709053,
  "rsi_1m_zone": "neutral",
  "atr_bucket": "mid",
  "vol_bucket": "mid",
  "lreg": "bear",
  "mreg": "bear",
  "htf1h": "bull",
  "ret_5": 0.29253419614847515,
  "ret_20": -0.10460349598912234,
  "slope_10": -28.65109090909069,
  "meta16": "fresh",
  "meta24": "gt4",
  "meta26": "young",
  "meta34": "weak",
  "meta40": "240-480",
  "meta41": "ABC_under_D",
  "meta46": "high",
  "meta87": "deep_low",
  "meta88": "bear",
  "meta89": "mixed",
  "meta94": "D>C>A>B",
  "meta96": "strong_up"
}
02 / trade record

The candidate becomes a resolved historical trade.

{
  "record_type": "trade",
  "timestamp": "2026-08-25 16:39:00+00:00",
  "timeframe": "5m",
  "side": "short",
  "symbol": "BTCUSDT",
  "event": "rb_cross_dn",
  "bar_idx": 3606,
  "tf_idx": 721,
  "rb_pos": "high",
  "rb_zone": "upper",
  "band_order": "B>A>D>C",
  "rb_hier": "cross_dn",
  "rb_reg": "bull",
  "d_slope_bucket": "strong_down",
  "rsi": 50.05958456607261,
  "rsi_1m": 47.53570908709053,
  "lreg": "bear",
  "mreg": "bear",
  "htf1h": "bull",
  "label": 1,
  "pnl": 0.8,
  "entry_time": "2026-08-25 16:39:00+00:00",
  "exit_time": "2026-08-25 20:39:00+00:00",
  "label_ret20_up": 1,
  "label_horizon": 20
}
03 / execution

Same decision, execution lifecycle.

{
  "stage": "trade_open",
  "decision_id": "6172791ac27a9332",
  "bar_idx": 3606,
  "tf": "5m",
  "side": "short",
  "entry_px": 79218.34
}

{
  "stage": "trade_close",
  "decision_id": "6172791ac27a9332",
  "exit_ts": "2026-08-25 20:39:00+00:00",
  "exit_px": 78584.59328,
  "profit_pct": 0.8,
  "hold_bars": 240,
  "exit_reason": "tp"
}
04 / outcome availability

Future measurements have their own clock.

{
  "stage": "trade_mfe",
  "decision_id": "6172791ac27a9332",
  "bar_idx": 3606,
  "tf": "5m",
  "side": "short",
  "entry_ts": "2026-08-25 16:39:00+00:00",
  "window_frame": "1m",
  "entry_bar_excluded": true,
  "fixed_window": {
    "5": {
      "mfe_pct": 0.0282,
      "mae_pct": 0.14091,
      "label_available_ts": "2026-08-25 16:44:00+00:00"
    },
    "20": {
      "mfe_pct": 0.15936,
      "mae_pct": 0.19407,
      "label_available_ts": "2026-08-25 16:59:00+00:00"
    },
    "50": {
      "mfe_pct": 0.5886,
      "mae_pct": 0.19407,
      "label_available_ts": "2026-08-25 17:29:00+00:00"
    },
    "240": {
      "mfe_pct": 0.84367,
      "mae_pct": 0.20407,
      "label_available_ts": "2026-08-25 20:39:00+00:00"
    }
  }
}
One observation, different moments in time. The decision exists at 16:39 UTC. A 5-minute outcome can only exist later. A 20-minute outcome later still. The 240-minute measurement is not available until 20:39 UTC. Vortraq keeps those distinctions explicit instead of treating a completed historical dataset as though all of its information existed at once.

The raw logs shown to users contain additional fields. This website example is shortened for readability, not because the full audit is hidden from the user.

Validation

Backtests test the strategy. Reality tests the backtester.

The conventional validation stack is already here

Parity is an extra test, not a replacement for the usual ones.

Vortraq does not ask you to choose between established historical validation and live/offline parity. Use the familiar tests when they fit the question: walk-forward analysis, frozen holdouts and out-of-sample periods, Combinatorial Purged Cross-Validation (CPCV), purge and embargo controls, time-shift tests, permutation tests and Monte Carlo robustness analysis.

Those tests attack historical robustness, chronology, leakage risk and sensitivity from different angles. Parity sits beside them and tests something they generally cannot: whether the historical research path can reproduce the state that the live system actually observed.

WALK-FORWARD
HOLDOUT / OOS
CPCV
PURGE + EMBARGO
TIME-SHIFT
PERMUTATION
MONTE CARLO
+
LIVE / OFFLINE PARITY

Historical validation asks whether the edge survives unseen time. Live/offline parity asks a different question: can a separate offline backtest, run after the live session over the same historical candles, independently reconstruct the information that actually existed when each live decision was made?

Parity validates the decision engine, not the exchange. The live run and the parity backtest are not one loop checking itself. Vortraq first records what the live engine actually saw. Afterwards, a separate offline backtest re-fetches the historical candles from the canonical signal-data source rather than replaying the live recording, then rebuilds the state through the research path as a later independent run. The two outputs are then compared. Fill behaviour, latency and venue execution are separate questions and remain recorded as such.

The published sessions are examples, not a certificate. The point is that parity can be tested again on your own run. Record the live decision-time state, run a fresh offline backtest across the same matched candles, and compare the independently produced state and decisions field by field.

Short sessions have limits. A long-lookback field may barely move during a 27-hour run, so matching it is weaker evidence than repeatedly exercising a fast-changing field. Longer runs, restarts, interruptions and different market conditions stress different parts of state reconstruction. The published sessions show the mechanism being measured; they are not presented as the end of validation.

Uninterrupted live session
130 / 130

evaluated strategy & feature fields matched

27 hours · 1,622 live bars.

Network interruption test
100%

comparable completed fields and rows after recovery

24-hour session · deliberate 10–15 minute blackout.

Why field-level parity matters

PnL can agree while an internal feature is already wrong.

A stale RSI, wrong volume bucket or misaligned higher-timeframe state may not change the next trade. An equity curve can still look fine while the defect waits for the market condition that exposes it.

Vortraq compares the recorded live decision state with the same instant produced later by a separate historical backtest — candles, derived features, structural state and subsequent decision. The offline side is rebuilt from the historical input; it is not a copy of the live engine's calculated state. Similar PnL is therefore not used as a substitute for parity.

LIVE RUN
RECORDED LIVE STATE
RE-FETCH HISTORICAL CANDLES
SEPARATE OFFLINE BACKTEST
INDEPENDENTLY REBUILT STATE
FIELD / DECISION PARITY
A second thing this test observes

The historical source has to agree with what arrived live.

The offline run does not replay the candle record captured by the live session. It obtains the matching historical candles again from the canonical signal-data source. When the independently rebuilt state then matches the recorded live state, that also provides evidence that the later historical candle data agreed sufficiently with the market data observed live across that tested window to reproduce the compared state.

This is evidence for the tested period and compared outputs, not a claim that a data provider's live and historical feeds can never diverge.

HISTORICAL VALIDATION

Does the edge survive when you change the test?

Walk-forward, holdout/OOS, CPCV with purge and embargo, time-shift, permutation and Monte Carlo attack robustness, chronology and leakage risk before parity is even considered.

STATE PARITY

Did offline know only what live knew?

Reality provides the independent timestamped evidence for the research machinery.

EXECUTION / VENUE

Did venue mechanics change the result?

Hyperliquid Lab studies one-position behaviour; live fills and exchange records remain the stronger evidence for what the venue actually realised.

What parity does not prove

Agreement is not independent proof that shared logic is correct.

Live and offline processing share parts of Vortraq's feature and structural computation. Parity therefore tests whether state was available and reproducibly reconstructed across the two operating paths; it does not independently prove that every shared definition or timing assumption is itself correct.

A causal or timing assumption implemented identically in both paths could reproduce perfectly and still be wrong. Leakage tests, chronology checks, label-availability enforcement and other causal validation address that different question.

Don't take our parity claim on faith. Re-run it.

TRIAL

See the evidence before you pay.

The trial is intended to expose enough of the underlying state/parity machinery to let a serious user test the claim rather than trust the website.

PRO

Keep the evidence.

Pro keeps the deeper local records available as part of the research surface, so the corresponding period can be run again as a separate offline backtest and challenged against the original live evidence later.

These are measured validation sessions, not a claim that software, networks or exchanges can never fail.

Technical paper

When Backtests Agree but the Data Doesn't

Why conventional validation cannot by itself prove that a trading model learned from the state a live system actually saw — and why Vortraq measures live/offline observational parity directly.

Method

The tests are public. The features are not.

Vortraq's structural engines and feature definitions are part of the product and stay closed. The validation machinery around them is not the edge, and hiding it makes the research harder to evaluate.

Baseline validation

The familiar tests are not missing because we talk more about parity.

Vortraq includes the standard validation methods a quantitative researcher would expect to reach for: walk-forward analysis, holdout/OOS testing, CPCV, purging and embargoes, time-shift tests, permutation tests and Monte Carlo robustness analysis. They are configurable research tools, not a single compulsory ceremony.

The additional machinery below exists because passing familiar tests still does not answer every question about causality, label availability or whether live and historical state were actually the same.

Walk-forward analysis

Train on what was available, move forward in time, retrain at the cadence the experiment specifies and evaluate the next unseen period. Vortraq treats retraining cadence and history length as part of the question, not decorative settings.

Holdout / out-of-sample

Freeze the chosen process and test it on data excluded from fitting or selection. A long frozen holdout and a frequently retrained walk-forward answer different questions, so Vortraq exposes both rather than pretending they are interchangeable.

Monte Carlo robustness

Resample or perturb the realised outcome sequence to examine sensitivity, drawdown and path dependence. Monte Carlo is useful robustness evidence, but it is not treated as proof that the underlying observations were causal or correctly reconstructed.

Purged combinatorial cross-validation

Fold boundaries are purged with an embargo so overlapping trades cannot quietly straddle training and test. The point is simple: a chronological split is not clean merely because the timestamps are different.

Label availability

A training row is admitted only when its outcome was actually knowable at the training cut-off, not merely because the trade had already opened. Future outcomes do not become historical data early just because the row exists.

Causal thresholds

Where a target needs a data-derived threshold, that threshold is derived from the training side of the experiment rather than from the full future-inclusive population.

Outcome-feature refusal

The selectable feature universe is screened for outcome and forward-path fields. Forbidden outcome information is refused rather than accepted with a warning.

Reverse-label tests

Flip the labels and ask whether the machinery follows them. If a model trained on the opposite outcome still appears to find the original winner, that is evidence of artefact rather than something to celebrate.

Time-shift tests

Deliberately decouple decision-time features from their outcomes. A relationship that remains suspiciously intact after its timing has been broken deserves investigation, not a higher confidence threshold.

Permutation tests

Shuffled-label runs show what the same machinery can manufacture from noise. They are a screen against self-deception, not a certificate that the surviving strategy will make money.

No single validation is the verdict

Walk-forward, holdouts, CPCV, Monte Carlo and diagnostic tests answer different questions. None of them — individually or together — turns a historical result into a promise about tomorrow.

Why publish this? Because validation behaviour is not the proprietary strategy logic. You should be able to see how Vortraq tries to catch leakage without being given the structural recipe it is protecting.

Multi-timeframe research

Why look from the bottom up?

Higher timeframes have obvious attractions: fewer trades, larger price targets and proportionally less sensitivity to execution costs.

But higher-timeframe structure does not begin when a higher-timeframe candle closes. It is assembled from the lower-timeframe behaviour inside it. By the time that structure becomes obvious at a higher resolution, much of the activity that created it has already occurred.

Lower timeframes therefore provide Vortraq with more than trading opportunities. They provide greater observational resolution into how market structure develops, while also producing the larger sample populations needed to test machine-learning relationships with less dependence on small-sample luck.

Vortraq does not treat this as an argument against higher timeframes. It trades across 1m, 5m, 15m, 1h and 4h, and information can complement research across timeframe boundaries.

The question is not simply which timeframe performs best, but what the development of structure at one resolution can tell us about the market observed at another.

Vortraq technical research note

When Backtests Agree but the Data Doesn't

Why validation is not the same as proving that a trading system learned from reality.

Machine learningLive / offline parityResearch integrityVortraq Ltd · 2026
Scope. This paper describes the problem Vortraq measures, the evidence produced by its validation process and the limits of that evidence. It intentionally does not disclose the proprietary mechanisms used to establish observation identity, reconstruct state, associate outcomes or enforce causal eligibility.
About this paper. This paper was produced by Vortraq Ltd. It describes Vortraq's engineering methodology and internal testing; it is not independent academic research or third-party validation. The stated test results apply only to the stated test conditions and do not demonstrate or imply future trading performance. Nothing in this paper is investment advice or a recommendation to acquire, dispose of or hold any cryptoasset or financial instrument.

Machine learning in trading is usually judged by its results: accuracy, AUC, profit and loss, Sharpe ratio, walk-forward performance, out-of-sample performance, Monte Carlo robustness and stability across market regimes.

These are useful measurements. But they all come after a more fundamental question:

Was the model actually shown an accurate representation of what the trading system could have known at the moment each decision was made?

If the answer is uncertain, everything measured afterwards inherits that uncertainty.

This is a problem Vortraq encountered directly during its development. It changed how we think about machine learning, backtesting and evidence.

The problem comes before the model

Consider a machine-learning trading observation. At 13:06, the system sees a collection of market information: price and recent returns, indicator values, structural market state, higher-timeframe context, volatility, previous events and strategy state. Call that collection X(13:06).

The model makes a decision using that state. Later, once enough market movement has occurred, an outcome Y can be associated with the observation. Training therefore attempts to learn:

X(t) → Y(t+h)

In a completed historical dataset, however, the system already possesses information about everything that happened afterwards. The research environment must reconstruct the past while behaving as though the future has not happened yet.

That is considerably harder than simply processing historical candles in chronological order.

  • A feature can be associated with the wrong observation.
  • A derived state can become available earlier in research than it could have been live.
  • An outcome can exist in the dataset before the simulated system should be allowed to learn from it.
  • Different stages of an experiment can unknowingly operate on different populations.
  • A reconstructed historical state can differ subtly from the state originally observed live.

None of these necessarily produces an obvious software failure. The program still runs. The model still trains. The backtest still produces trades. And the results can look excellent.

Machine learning cannot tell you that its teacher is wrong

A machine-learning algorithm has no independent understanding of the market it is being shown. It learns relationships in its training data.

If the data contain a genuine relationship, the model can attempt to learn it. If the data contain noise, the model can learn around that noise. If the data contain a systematic error, the model can learn the error. And if information unavailable at decision time has entered the observations, a sufficiently capable model may become exceptionally good at exploiting it.

Better machine-learning performance does not prove better trading information.

There is another, less obvious failure mode. Suppose most historical observations are correct, but a small proportion are not. The model may encounter what appear to be contradictory examples:

X → WIN     X → LOSS     X → WIN

If some of those contradictions were introduced by reconstruction, association or timing errors rather than by the market itself, the model cannot know that. To the model, both market uncertainty and data-integrity errors look like data.

The consequences can appear as weak confidence separation, unstable feature importance, poor generalisation or predictive performance collapsing toward randomness. Adding more data does not necessarily solve the problem. If the process producing the data remains imperfect, more data can simply mean more contradictory evidence.

Before asking how well the model can learn, ask what exactly you are teaching it.

Why conventional validation is not enough

Vortraq uses conventional quantitative validation techniques. They remain valuable. But during development we learned an important distinction: a validation test can test only the failure modes it was designed to detect. Passing it does not prove that every other failure mode is absent.

Out-of-sample testing

Separating training and evaluation periods reduces one of the most obvious forms of overfitting. But chronological separation does not establish that the observations inside either period were causally correct. A perfectly separated test set can still contain incorrectly reconstructed information.

Walk-forward testing

Walk-forward evaluation improves realism by repeatedly training on past information and evaluating on later periods. It is an important test. But it still assumes that the state presented at each simulated historical moment accurately represents what could have existed at that moment. Walk-forward validation can therefore evaluate an invalid historical representation extremely rigorously.

Monte Carlo analysis

Monte Carlo techniques can provide valuable information about sequence risk, drawdown behaviour and sensitivity to observed outcomes. They do not establish whether the observations that produced those outcomes were temporally and structurally valid in the first place. Statistical robustness and data integrity are different questions.

Leakage diagnostics

Explicit leakage tests are essential. But they too have limits. A leakage test searches for a particular signature of leakage. If contamination enters through a mechanism the test does not perturb or observe, the test can pass.

Vortraq encountered this problem during development. A significant temporal-integrity issue existed while dedicated diagnostic testing — including a temporal-shift test — reported no leakage. The diagnostic was not dishonest. It simply was not capable of detecting the particular mechanism responsible.

That experience changed our standard of evidence.

Stop asking whether we can detect a difference

Most validation approaches effectively ask: Can we find evidence that the historical system differs from reality?

Vortraq eventually began asking a different question:

Can the historical system reproduce the state that reality actually produced?

Consider a live trading system at 13:06. It observes Xlive(13:06). Later, after that period has become historical, the research system reconstructs the same moment as Xoffline(13:06).

The question is no longer whether their resulting equity curves look similar.

X_live(13:06) =? X_offline(13:06)

And then the same question at 13:07, 13:08, 13:09 and onward across the tested session. We refer to this as observational parity.

Same trades do not prove the same observations

Output agreement can hide internal disagreement.

StateLiveOffline
Feature A0.820.77
Feature B0.310.31
Market regimeBullBull
DecisionLongLong

Both systems make the same trade. Their fills may be similar. Their resulting P&L may be nearly identical. A live-versus-backtest equity comparison therefore looks excellent.

But parity has already failed. The offline research environment did not reconstruct the information seen live. The error may simply not have affected this particular decision.

That distinction becomes especially important in machine learning. A feature that has little influence on today's model may become important to a model trained tomorrow. Two incorrect features can even offset one another and produce the correct final decision for the wrong reasons.

D_live = D_offline    does not establish    X_live = X_offline

This is why Vortraq does not treat similar trades or similar P&L as proof of research/live parity.

Same code does not prove the same world

Running identical strategy code in research and live trading is good engineering practice. Using the same execution engine is better still. But neither proves observational parity.

If the same function runs in both environments but the information supplied to it differs, identical code can legitimately produce different internal states or decisions. Even if it happens to produce the same decision, the underlying research representation remains different.

The stronger question is not merely “Did we execute the same code?” It is “Did the same code see the same world?”

Vortraq's approach

Vortraq records decision-time state during live operation. Afterward, the corresponding historical candles are fetched again from the canonical signal-data source rather than replayed from the live recording, and the period is run through a separate offline backtest using the historical research path. That later run independently reconstructs the historical observations, which are then compared directly with the recorded live state.

The comparison is performed on the state used by the system rather than relying solely on resulting trades or profitability. This allows discrepancies to be detected even when they do not alter the final trading decision.

The internal mechanisms used to establish observation identity, reconstruct state, associate outcomes and preserve causal eligibility are part of Vortraq's proprietary architecture and are intentionally not described in this paper.

A boundary of the test: the live and offline paths share parts of Vortraq's feature and structural computation. Observational parity therefore does not independently prove that shared logic is conceptually correct. If the same causal or timing assumption were wrong in both paths, they could agree perfectly. Parity tests reproducibility and availability across the operating paths; separate leakage, chronology and causal-eligibility tests address whether the shared research assumptions themselves are valid.

The objective of the test is public. The implementation is not.

What we tested

Vortraq has been subjected to extended live-versus-offline parity testing in which live decision-time observations were recorded first, then the matching period was run separately through the offline historical backtest and the independently reconstructed observations were compared.

Testing has included extended continuous market operation, large collections of simultaneously changing state fields, multiple timeframe-derived states, structural and indicator information, decision-related state, interruption and recovery scenarios, and repeated historical reconstruction.

In one documented session, 130 compared state fields reproduced without mismatch across the tested observations during approximately 27 hours of live operation.

A subsequent extended test ran for approximately 42 hours and produced byte-identical compared output across the tested live and reconstructed paths.

What “byte-identical” means here. It does not mean that every byte produced anywhere within Vortraq is guaranteed identical under every possible condition. It means that the specified outputs included in that experiment were compared across the tested live and offline processing paths and were identical for that experiment. Reproducibility claims should describe exactly what was measured.

What this proves — and what it does not

Observational parity does not prove that a trading strategy will be profitable. It does not prove that a machine-learning model contains predictive information. It does not eliminate market uncertainty. It does not prove that software is incapable of containing another defect. It does not guarantee that every future market condition will reproduce perfectly.

What it provides is evidence for a narrower proposition:

Across the tested period, fields and operating conditions, the historical research process reproduced the decision-time state recorded by the live system.

That proposition matters because virtually every subsequent research conclusion depends upon it.

Why this matters for machine learning

A trading model is trained on historical observations because we hope those observations represent situations the model could encounter live. If that assumption is wrong, model statistics become difficult to interpret.

Consider an AUC of 0.60. Is that genuine predictive information, information accidentally unavailable at decision time, an association error, or a population inconsistency? AUC alone cannot answer.

Now consider an AUC of 0.50. Does the market contain no learnable relationship? Are the features poor? Or has reconstruction error introduced enough contradictory information to obscure an otherwise weak relationship? Again, the model cannot answer.

This is why Vortraq treats observational integrity as a prerequisite for interpreting ML performance rather than as another performance metric.

Establish observation integrity → then measure what the model can learn

Learning from outcomes creates another clock

There is also an important difference between information used to make a prediction and information used later to teach the model.

At time t, a model may make a prediction. The outcome used to evaluate that prediction may not become knowable until t+h. The fact that the outcome already exists somewhere inside a completed historical dataset does not mean a simulated trading system should be permitted to learn from it at time t.

Historical availability and simulated causal availability are not the same thing.

An adaptive machine-learning system therefore has to reason not only about when a prediction was made, but also when its outcome could legitimately have become known. Vortraq treats this distinction as part of the integrity of the learning process. The internal mechanisms enforcing it remain proprietary.

Observation identity matters

A prediction, trade and eventual outcome must refer to the correct original observation. Matching records because they appear close enough — same timestamp, same direction, same timeframe, similar state — is not necessarily sufficient in a complex adaptive system.

Once observations travel through multiple research, trading and learning stages, identity becomes part of correctness. Vortraq therefore treats the continuity of an observation through its lifecycle as a first-class system concern.

Again, this paper deliberately describes the requirement rather than the implementation. Knowing that identity matters is not the same thing as knowing how Vortraq establishes and protects it.

Why we don't use profitability as proof

A profitable backtest is evidence that a particular historical simulation produced profit. It is not proof that the simulation accurately represented the information available to a live system.

Likewise, a profitable live period does not prove that a research methodology is valid. Markets contain variance. Strategies encounter favourable and unfavourable regimes. Unknown edge and luck can both produce attractive short-term results.

Vortraq therefore separates two questions:

Does the strategy possess an edge?    /    Can the research machinery itself be trusted?

The first question is inherently uncertain. The second can be subjected to engineering measurement. Our parity work addresses the second.

Why we publish this

Vortraq is a commercial product. This paper should therefore not be interpreted as independent third-party research. It describes our engineering philosophy, development experience and internal test evidence.

We publish it because we believe users evaluating automated trading and machine-learning systems should ask questions beyond historical profitability.

Ask whether research and live trading use the same logic. Then go further.

  • Ask whether anyone has measured whether they actually produced the same state.
  • Ask what was compared.
  • Ask for mismatch counts.
  • Ask how long the test ran.
  • Ask whether changing internal values were examined or only final trades.
  • Ask whether interruptions and recovery were tested.
  • Ask whether a machine-learning outcome can enter training before it would have been knowable.
  • Ask whether an observation being trained later can be proven to be the same observation originally evaluated.

And distinguish carefully between architecture that should produce parity and evidence that parity actually occurred. They are not the same thing.

The principle

Machine learning does not know whether the world we give it is real. It only knows the data.

So before asking “How intelligent is the model?”, we believe a trading system should first be able to answer “How faithfully did we reconstruct what the model was allowed to know?”

That question does not produce an exciting equity curve. It does something more fundamental. It determines whether the equity curve, the model statistics and the research built on top of them deserve to be interpreted at all.

Before teaching the machine, prove the lesson is real.
Vortraq Ltd · Technical Research Note · 2026
This document describes engineering methodology and measured software-validation evidence. It is not investment advice, a profitability claim or a guarantee of future software behaviour.
Research stance

Evidence before belief.

Simple isn't disqualified.

An EMA cross doesn't become nonsense because it is familiar.

Simple doesn't mean useless.

Complex isn't automatically clever.

A custom indicator doesn't become useful because it has more maths behind it.

Complex doesn't mean useful.

Vortraq isn't anti-indicator. It is anti-untested certainty.

How much proof do you want?

That is the trader's decision. One person may deliberately trade a model that is strongest inside the macro environment they believe they are in. Another may insist that an edge survives years of regime changes, repeated retrains, holdouts and a drawdown ceiling.

Vortraq does not force the second trader's burden of proof onto the first one.

The depth is available. It is not compulsory homework.

What our own research taught us

A good-looking period is not the same thing as a stable edge.

The reason to test three months, three years, different halves of history, different retraining assumptions and eventually much longer spans is not because Vortraq requires endless research. It is because some traders want to know whether what they found survives the kinds of change they personally care about.

Environment

It lives on your screen. Make it comfortable.

This is not a major trading feature. It is simply useful. Vortraq can switch environment depending on where you're working — standard dark, a light scheme for bright conditions, or Matrix green when subtlety has left the building.

Light environment.
Matrix green.

The three buttons in the website header are interactive too. Because why not.

Pricing

Pick how far you want to take it.

Standard is not a crippled version of Vortraq. Pro is the full research laboratory, including the deeper evidence trail for users who want to inspect and reproduce what the system saw.

TRIAL
£0

7 days · no card required.

  • Explore the full research environment
  • Backtest / replay / validation tools
  • Hyperliquid Lab venue-behaviour replay
  • ML Workshop and deeper research surfaces
  • Inspect raw evidence and parity material during the trial
  • Live execution disabled
PRO
£599 /month

or £5,391 / yearsave £1,797

  • Everything in Standard
  • Broader policy / candidate systems
  • Hyperliquid Lab + research modes
  • Rainbow V2 + Stretch
  • Deeper local research / evidence loop
  • Retain raw snapshots, event state and parity evidence for forensic research
  • Re-run offline comparisons against your own recorded live data
  • Expanded licence / support terms
Why the split?

Standard can use Vortraq seriously without exposing the whole internal research substrate.

Pro is for the user who wants the laboratory records themselves — the detailed state, provenance and parity material that lets them inspect not just the result, but the machinery that produced it.

Launch list

How close do you want to get?

Choose the level of involvement you actually want. The limits below reflect how much onboarding and hands-on support can be done properly, not a countdown timer.

NO CAP

Just tell me when it's live.

No commitment. Launch and meaningful milestone updates only.

  • Notified before public launch
  • Launch pricing locked at signup
  • No constant sales emails
EARLY TESTER GROUP · 10

Help break it before launch.

The tester group is small because useful testing means reading the reports, reproducing failures and following up rather than collecting names.

For people willing to run pre-launch builds, find problems and give useful feedback.

  • 50% off any tier for 6 months at launch
  • Direct testing feedback line
  • Founding tester credit if you want it
  • First look at new work

We read every application. Selected testers will be contacted within 7 days.

We read every application. Reservation confirmations are sent by email.

Company & legal

Risk Disclosure

Trading cryptoassets and other financial instruments involves substantial risk. Losses can occur rapidly and historical, simulated, backtested, paper-trading or machine-learning results do not guarantee future performance.

Research results can differ from live outcomes because of market conditions, liquidity, spreads, fees, slippage, latency, exchange behaviour, connectivity, outages, order handling and other factors. Validation and parity testing provide evidence about specified software behaviour under specified test conditions; they do not establish future profitability.

Vortraq can automate rules selected or configured by the user. Automation does not remove trading risk. Users are responsible for deciding whether a strategy, level of risk and connected venue are appropriate for them.

Cryptoasset markets can be highly volatile. Users should not trade with money they cannot afford to lose.

Company & legal

Privacy

Data minimisation

Vortraq Ltd's approach is to collect and retain only information reasonably necessary to operate the website, provide licences and support, administer customer relationships, meet legal obligations and protect the service.

Exchange API credentials

Vortraq Ltd does not need customers to provide exchange API credentials to the company and does not keep customers' API credentials. Customers retain control of their exchange access and can manage or revoke that access through the relevant exchange.

Website and account information

The production privacy notice will identify the personal information actually collected by the launch website and service, the purposes and lawful bases for processing, processors used for payments/hosting/email or analytics, retention periods, international transfers where applicable, and customer data-protection rights.

Launch note. This section intentionally does not invent processors, retention periods or telemetry that have not yet been confirmed. It must be completed against the production stack before personal data is collected from the public.
Company & legal

Terms & Software Licence

Vortraq is licensed software. Purchase of a subscription does not transfer ownership of Vortraq, its source code, proprietary research architecture, internal ontology, models, documentation or other intellectual property.

Users are responsible for the strategies and settings they deploy, their connected exchange accounts, compliance with exchange terms and applicable law, and all trading decisions made through their accounts.

Vortraq depends in part on third-party exchanges, market-data sources, networks and APIs. Vortraq Ltd cannot guarantee continuous availability or unchanged behaviour of third-party services.

Backtests, simulations, validation reports, model statistics and technical research are research outputs and are not promises of future performance.

The final customer terms will also set out subscription and cancellation mechanics, permitted use, licence restrictions, support, updates, suspension/termination, warranties, liability limits, consumer rights where applicable and governing law.

Launch note. Final contractual terms should be reviewed against the actual checkout flow, subscription/refund policy and production licence mechanism before paid subscriptions are accepted.
FAQ

The questions we'd rather answer plainly.

No. It gives you trading engines, policies, analysis, ML and execution machinery. It does not promise one immutable profitable strategy.
No. Connect a supported account, choose a signal and timeframe, and Vortraq can trade without a model. ML is an optional layer.
Yes. Trade Policy, structural filters, regimes, channels and Trade Analysis can all lead to deterministic logic without ML.
It interrogates completed populations from many directions: structure, clocks, regimes, interactions, execution behaviour, policy diagnostics, stability and OOS work. Finding a relationship is the start of the investigation, not the result. Change its definition, period or surrounding conditions and see whether it survives. Vortraq gives you tools to expose fragility; it cannot stop a researcher choosing to ignore it.
No research platform can stop someone repeatedly testing ideas until something looks good. Vortraq does not pretend otherwise. Its job is to make the first attractive result easier to challenge: nearby definitions, different periods and regimes, chronological tests, alternative populations and OOS behaviour can all be examined before you decide what deserves belief. If a result only survives the exact question that discovered it, that is evidence too.
Yes, because they may be testing different things. History length, retraining cadence, holdout structure, embargo and validation mode are configurable on purpose. A model retrained weekly across years of history is not the same experiment as one frozen for a long holdout block. Vortraq exposes those choices rather than collapsing them into one universal “validated” badge.
No. Parity is an additional validation layer. Vortraq also provides walk-forward analysis, holdout and out-of-sample testing, CPCV with purging and embargoes, time-shift tests, permutation tests and Monte Carlo robustness analysis. Use the methods that fit the research question; parity then addresses the separate question of whether the historical engine reconstructed what live actually saw.
New trade admission can fail closed while missing closed candles are restored and continuity is rebuilt. Existing position management is treated separately.
It is something Vortraq measures rather than assumes. The comparison is not a live loop checking itself: Vortraq records the live decision-time state, then later re-fetches the matching historical candles from the canonical signal-data source and runs a separate offline backtest that independently rebuilds the state for comparison. Recent controlled tests include complete parity across the evaluated fields in both uninterrupted and deliberate network-loss sessions.
No. Vortraq Ltd does not need customers to provide exchange API credentials to the company and does not keep customers' API credentials. Exchange access remains under the customer's control. Vortraq does not take custody of customer funds.
Because the position mechanics differ. MEXC can keep opposing positions open in the supported setup; Hyperliquid nets exposure, so opposite signals can reduce or reverse what is already open.
It takes an existing raw or Trade Policy population and replays it under Hyperliquid-style one-net-position mechanics. It does not create a trading population or teach ML. It is a venue-fit and research view: the same signals can behave differently when opposing trades reduce, flatten or reverse existing exposure.
Yes. The trial is intended to expose the evidence deeply enough to test the claim. Pro keeps the detailed local records so you can take a recorded live period, run a separate offline backtest over the matching historical candles, and compare the independently produced state again later.
Standard still shows trading history, health, validation and parity results. Pro keeps the much deeper snapshot/event/provenance material because that data is itself part of the research laboratory and can expose a large amount of Vortraq's internal ontology.
Risk and disclosures
Vortraq is trading software, not advice.

Vortraq is a research, validation and execution environment for algorithmic trading. It is not investment advice. Vortraq's makers are not licensed investment advisers and nothing on this site constitutes a recommendation to buy, sell or hold any asset.

No guarantee of profitability.

Trading involves substantial risk of loss and is not suitable for every investor. Any performance figures, validation results or example outputs shown on this site are illustrative and do not guarantee future results. You may lose some or all of the capital you commit.

Refunds.

A 7-day free trial is available before any paid subscription begins. Once a paid subscription starts, no refunds are issued for the current billing period. You can cancel at any time and retain access until the end of the period already paid for.

Your responsibility.

You are responsible for the trading decisions you take, the strategies you deploy and the venues you connect Vortraq to. Backtests and validation tests are research tools, not predictions. Vortraq is software; it does not eliminate risk, it exposes it.

API credentials and custody.

Vortraq Ltd does not need customers to provide exchange API credentials to the company and does not keep customers' API credentials. Exchange access remains under the customer's control. Vortraq does not take custody of customer funds.

Vortraq Ltd · Registered in England and Wales · Company No. 17421751
Created by Tadas Siksnius
Registered office: 71–75 Shelton Street, Covent Garden, London, WC2H 9JQ, United Kingdom
© 2026 Vortraq Ltd · · · ·