AlphaRatings come from public on-chain and venue data and may be incomplete. We see on-chain only — exposure on centralised exchanges is invisible to us.Methodology
ASPERN
⌘K

Methodology

Version 0.2 · every figure on this site links back to the version that produced it.

What we measure

Return over observed history, worst peak-to-trough drawdown, downside-adjusted risk-return (Sortino), the length of live history, and a confidence figure derived from how much history and how many observations support the rest. Confidence is reported separately and never folded into a score: thirty days of data must not look like three years.

Provenance grades

A
— attested through ASPERN: pre-trade committed, sequence complete, anchored on-chain.
A−
— pre-registered with us before going live.
B
— verifiable on-chain or at the venue, but the operator’s full cohort is unproven.
C
— platform-reported, not independently verifiable.
D
— claimed only.

Everything currently on this site is grade B. Nothing is A until the attestation SDK ships.

How we decide a vault is agent-managed

No registry of agent-managed vaults exists, and self-declaration is unreliable in both directions — marketing-led operators overclaim autonomy, serious quantitative teams underclaim it. So we infer, and we publish the evidence and the confidence alongside every claim. Version 0.1 uses platform membership and name signals only. Behavioural inference is next: rebalance clock regularity, activity uniformity across the 24-hour cycle, reaction latency to market events, and the entropy of trade sizing.

A standing rule: inference may raise a flag, but only a declaration or an attestation by the operator can lower a rating.

How we decide a vault is run by software

Three sources of evidence, ranked by how hard each is to fake, and every claim carries the evidence behind it.

Platform membership
— the venue exists to run autonomous strategies. Strong, and maintained by someone other than us.
Execution timing
— trading continuously through every six-hour window including overnight, evenly spread across the day, at rates and regularity beyond sustained manual trading, with fills landing on clock boundaries. A name is a marketing choice; timing is not.
Name signals
— the weakest, and treated as such. It is enough to raise a flag, never enough to be confident.

The overnight signal is undefined below a day of observation, and its weight is redistributed rather than counted as zero: absence of evidence is not evidence of a human. Confidence grows with sample size and observation span, never with the verdict. And the standing rule — inference may raise a flag, but only a declaration or an attestation by the operator can lower a rating.

Capacity

How much capital a strategy could absorb before its own market impact erodes the edge. Two limits are computed and the tighter one wins: the ratio of realised edge to realised cost, and a 10% ceiling on the share of venue daily volume any strategy is assumed able to take. Both describe a trade size, which is then scaled to capital using the strategy’s own ratio of capital to trade size.

Three bounds keep the result honest. A measured per-trade edge above 100bps is clipped, because at that level it is a directional move that happened to go the right way rather than repeatable alpha, and squaring it would turn one lucky month into a billion-dollar claim. The result is capped at 50× current capital, beyond which it is unfalsifiable. And where either cap binds, the figure is published as a floor rather than an estimate.

Where a return came from

A headline return is a claim about skill; what it usually contains is a market move minus costs nobody itemised. We decompose it into gross realised profit, trading fees, funding paid or received, and unrealised marks kept separate because they are not money until closed. Fees are exact per fill and funding exact per position.

Slippage is not separated. Without a book snapshot at the moment of each fill it cannot be measured, so it remains inside gross and is declared rather than estimated. An invented slippage line would be worse than an honest omission.

Crowding

Exposure is compared across vaults on signed positions, so hedged opposites score as opposites rather than as similar. Instruments are flagged only when both conditions hold at once: a material share of venue open interest, and heavily one-sided positioning. Either alone is survivable; together they unwind as a group.

No crowding figure is published without a coverage ratio against venue open interest. A cluster size with no denominator is a guess wearing a suit. Book clusters use single-linkage grouping, so a cluster’s average similarity can sit below the joining threshold — transparent and arguable, which we prefer to a clever method nobody can interrogate.

Known limitations

  • We see on-chain only. Exposure held on centralised exchanges is invisible to us and always will be without operator cooperation.
  • Venue history is downsampled. Hyperliquid returns a limited number of points per vault, so early history is coarse. Our own daily snapshots accumulate from the day we started observing.
  • Capacity is inferred, not measured. It comes from a strategy’s own realised cost rather than from a book, and is often reported as a floor.
  • No cohort yet. We cannot yet tell you how many other strategies an operator ran and quietly abandoned.
  • Slippage is not separated. Fees and funding are exact; slippage stays inside gross because it cannot be measured without a book snapshot.
  • Depositor lists are capped at 100 by the venue. A wallet count of 100 is a floor, and concentration is computed over what we can see.

Independence

Ratings are free and public. There is no paid placement, no sponsored ranking and no advertising. We do not trade on our own data.

Methodology — ASPERN