As of: 2026-08-21 · Patch 3.1 · 58 agents ranked
This document describes fully and honestly how every agent on this page receives its score, its position and its verdict (UNDERRATED / OVERRATED / ELITE / STANDARD / NEW) — including every formula, every weight and every known weakness. None of it is secret; all of it also lives on the page itself.
We do not buy opinions, we collect them — and measure part of it ourselves.
| Source | Type | Notable |
|---|---|---|
| game8.gg | Major guide | Leads the board, also provides builds, teams, engines, discs |
| prydwen.gg | Curated guide | Maintained monthly, also rates Defense/Anomaly roles |
| icy-veins.com | Guide collection | Editorial tier list |
| u7buy.com | Editorial blog | Independent second opinion |
| mobalytics.gg | Theorycrafting platform | Idiosyncratic, data-driven voice |
Every source is fetched fresh on every refresh (6-hour cycle + manual button). When sources drop out (Cloudflare, redesign), we keep computing with the remaining ones — confidence (see 4.2) then automatically decreases.
| Source | What we measure from it |
|---|---|
| nanoka.cc (datamine) | Exact skill multipliers (Ultimate, EX, Basic, Daze) per agent — the basis for damage and stun |
| game8 build pages (57 agents) | All recommended teams per agent — the basis for team fit |
| Official news (HoYoverse, HoYoLAB, Gematsu) | Version, banners, events — controls which agents may be ranked at all (beta agents without gameplay data get no verdict) |
Each voice slots agents into grades (T0 through T4). We translate every grade into points:
T0 = 95 T0.5 = 88 T1+ = 82 T1 = 75 T1.5 = 68 T2 = 60 T2.5 = 52 T3 = 42 T4 = 30
For every agent we collect the point values of all sources that rank it and take the median (the middle value), not the average.
Why median: a single outlier voice (e.g. u7buy says T0.5 while everyone else says T2) must not pull a rating up. The median ignores outliers as long as the majority agrees.
Besides the median point value we compute for every agent the consensus rank: the position it would get if we sorted purely by the guide median — i.e. "where would the guides place this agent?". On point ties (many agents share the same median value) each gets the average rank of the group, so that sort order never arbitrarily decides verdicts.
The consensus rank is pure context info ("where would the guides put them?"). The yardstick for UNDERRATED/OVERRATED since the engine refactor is delta_own (see 5) — the distance of our measurements to the role midpoint, not the rank distance.
This is the part no guide can simply copy — here the page earns its own opinion.
From the nanoka datamine we read the exact skill multipliers per agent:
raw = Basic×0.25 + EX×0.40 + Chain×0.25 + Ult×0.10×decibel factor + Assist×0.10
Rotation model with decibel reality (the team shares 3000 DB): the Ultimate only counts at 10% — and full decibels only for Attack and Anomaly carries (they ult themselves), all other roles only ×0.2. Every category is captured as a chain sum: all hits of the combo, but per hit slot only the strongest variant (uncharged/charged/level-1/2/3 and parry variants are alternatives — exactly one runs in combat). Multiple basic chains (Nicole, Soldier 11) = combo variants → the strongest counts. Stance skills (Yanagi) count as basics. Defensive assist: highest parry tier (Light/Heavy/Chain) × 0.5 — reactive, situational. The raw value is then converted to a percentile within the role group (see 4.1): 90 means "deals more damage than 90% of the agents in its role".
The same datamine provides the daze multipliers. Again: raw value → percentile within the role group. For stun agents this is the core value; for other roles it barely matters.
We count in how many of the best teams an agent appears — but weighted:
1. Every team on a build page counts with the strength of the page owner (its median tier). A team on Miyabi's page is worth more than the same team on a T3 page. 2. Budget/F2P comps only count 25% — when an "Anby Budget" comp names Anby, that is placeholder advice, not a meta statement. 3. Bayes shrinkage: an agent appearing in only 2 teams so far (new agents!) is pulled cautiously toward the mean instead of punished: syn = (raw × teams + 50 × 6) / (teams + 6). 4. No cap: earlier the guide consensus used to cap our own team measurement here (max 55 at a T2.5 rating) — a circle that prevented UNDERRATED detection. Removed; budget teams still count only as a strength signal.
Anomaly agents (Burn, Shock, DoT) get their own group, because the damage formula cannot measure their damage fairly: a large part of their damage runs through time-based effects that are not contained in hit multipliers. That is why opinion and team count clearly more for them.
Not to be confused with strength! Confidence says: how sure the system is about THIS agent.
≥2 voices: conf = 1 − (mean deviation of all voices from the median / 45) 1 voice: conf = 0.60 + 0.05 × (number of available voices − 1), capped 0.60–0.85 0 voices: conf = 0.70 (pure own assessment) Penalty: agent in ≤2 teams → conf −0.05
If all 6 sources agree → ~100%. If two sources pull in opposite directions → visibly yellow/red. Confidence directly damps the opinion factor in the score formula — divided sources must not fully carry a rating.
These weights are deliberately fixed — they are our theory of what a role must deliver. They are NOT automatically fitted to the guides (that would be circular: opinion would confirm itself).
| Role | Opinion | Damage | Stun | Team |
|---|---|---|---|---|
| Attack (DPS) | 35% | 30% | 5% | 30% |
| Stun | 35% | 10% | 35% | 20% |
| Support | 30% | 5% | 10% | 55% |
| Anomaly | 40% | 10% | 5% | 45% |
| Defense | 30% | 5% | 20% | 45% |
How to read it: a supporter is defined by team fit (55%), a DPS by damage and opinion, a stunner by stun. A Pearson correlation between our measurements and the guide consensus runs as pure diagnostics (see META LAB → CALIBRATION) — it shows where our numbers and the guides diverge, but controls nothing.
score_raw = Σ( factor value × weight ) / Σ( weights ) score = score_raw × conf + 50.0 × (1 − conf)
Confidence no longer damps the opinion factor in the denominator (that let the relative weight of raw data RISE on source disagreement — mathematically backwards). Instead: global Bayes shrinkage toward the neutral middle 50 — everything uncertain = cautious statement. If a factor is completely missing (e.g. no voice), its weight is automatically redistributed to the rest — no hole appears.
Daily fluctuations (one source changes a grade today) must not throw the list around:
today's score = 30% yesterday's score + 70% fresh computation
Trend arrows (▲/▼) appear from ±3 positions of movement. On a patch change (new version) the history resets completely once — new meta, new merits.
Sorted by score; ties share the rank (1, 2, 3, 3, 5...). The DB bands:
S = top 8% (min 3) · A = top 30% · B = top 55% · C = top 80% · T = rest
Every verdict compares our measurements with the midpoint of its role (delta_own):
| Verdict | Condition | Meaning |
|---|---|---|
| ELITE (checked first) | score ≥78 AND top 12% | Top with us AND with the experts — the safe pick |
| UNDERRATED | delta_own ≥ +8.0 and score ≥50 | Better than its reputation — our measurements argue against the guides |
| OVERRATED | delta_own ≤ −8.0 | Weaker than its reputation — the guides overlook something |
| NEW | No source has ranked it | Too new — assessment comes from our numbers alone |
| STANDARD | Everything else | We largely agree with the experts |
delta_own = by how much our own measurements (damage, stun, team — role-weighted) raise/lower the agent versus the midpoint of its role (50 = role average of the normalized values). Scale-free — no longer a rank delta that produced false alarms in the mid-field (20 agents within 3 points) and mathematically never allowed UNDERRATED at the top (max 4 positions apart). No tier grouping: even a lone agent in its tier gets a true distance to the role midpoint.
Example: Nicole, supporter. Her measurements (team fit 92/100 — she appears in 33 of the best teams) lift her clearly above the supporter midpoint: delta_own ≈ +12 → UNDERRATED. No rank comparison, no tier view — just the question: "how far above the role average do her measured values lie?"
The ±8-point threshold exists because smaller distances are noise — only a real point distance between our measurements and the role average is a disagreement that deserves a name.
The composite is now a weighted geometric mean, fully transparent:
score = Π(factor^weight) (factors 5..100; missing factor = neutral 50)
Properties: all factors 50 → 50 (average), all 95 → 95 (peak); a weak factor drags more than in an arithmetic mean — a kit with a hole loses combat power, which is exactly the realism we want. No Bayes shrink toward the middle, no EMA smoothing, no clamp saturation: the number is 100% reproducible from the displayed factors and weights.
Class profiles are data-driven, not hand-picked. Every specialty (including future ones like Rupture, which arrived with 3.1) forms its own measurement peer group. The damage/stun weighting per class is derived from the class's median kit orientation, measured in GLOBAL percentiles: where does this class sit across all agents in damage output vs daze output? Stun classes legitimately weigh daze higher, attack classes damage — the game data itself defines the profile:
| Class | opinion | damage | stun | team |
|---|---|---|---|---|
| Attack | 0.30 | 0.224 | 0.196 | 0.28 |
| Stun | 0.30 | 0.157 | 0.263 | 0.28 |
| Support | 0.30 | 0.179 | 0.241 | 0.28 |
| Anomaly | 0.30 | 0.240 | 0.180 | 0.28 |
| Defense | 0.30 | 0.225 | 0.195 | 0.28 |
| Rupture | 0.30 | 0.166 | 0.254 | 0.28 |
(These numbers regenerate on every refresh — a new class automatically gets its own profile the moment ≥1 agent of it exists.)
Hollow Index vs pull economics: the index measures STRENGTH in a vacuum — pull decisions are investment economics (role, team cost, engine, opportunity cost). A high score is never a pull recommendation by itself; the SIGNALS view weighs both.
1. Transparency down to the formula. This document, the per-role weights (META LAB → CALIBRATION), the factor bars and the confidence bar on every card — all visible, nothing behind a blackbox. 2. Five independent voices instead of one. No single editor decides the opinion. Where the sources fight, you see it (confidence drops, the divergence table lists every deviation). 3. Own measurement instead of copy. Damage and stun come from exact datamine numbers, team fit from real build recommendations — not from retyping a tier list. 4. Admitted limits. Every WHY-box names the weaknesses of its own rating (measurable limits below). A system that knows its limits does not lie. 5. No fashion swings. EMA smoothing + patch reset: the list reacts to real meta changes, not to a source's daily mood. 6. Disagreement is allowed. UNDERRATED/OVERRATED openly show where we contradict the guides — with a per-agent justification. A page that only mirrors them would not need us.
follow-up mechanics are not contained in hit multipliers — anomaly agents therefore have their own weight profile, but the damage value remains a lower bound for them.
better in the damage value than it is in real runs.
shields/survival stay invisible, there is no shield data in the datamine.
estimates — the WHY box says so explicitly.
truth — exactly why we measure alongside.
6 tier lists ───► tier→points ───► MEDIAN ───► consensus rank (exp_rank) ┐
│
nanoka datamine ─► skill multipliers ─► dpr/daze percentiles ─────────────┤
├─► role weights
game8 builds ───► teams (tier-weighted, ─► syn (Bayes-smoothed) ──────────┤ (fixed)
budget teams count 25%) │
└─► SCORE
confidence (voice agreement, team coverage) ──► Bayes shrink →50 │
EMA (30% yesterday + 70% today) ──► smoothed score ◄────────────────────────┘
│
delta_own (measurements vs role midpoint) ──► UNDERRATED / OVERRATED /
ELITE / STANDARD / NEW
snapshots/<date>/) — every numberis traceable back to the source.
scripts/refresh.py, ~1600 lines, commented) lives in the repo.data stamp everywhere.