← BACK TO INDEX

HOLLOW INDEX — Methodology

As of: 2026-08-21 · Patch 3.1 · 58 agents ranked

This document describes fully and honestly how every agent on this page receives its score, its position and its verdict (UNDERRATED / OVERRATED / ELITE / STANDARD / NEW) — including every formula, every weight and every known weakness. None of it is secret; all of it also lives on the page itself.


1. The data sources

We do not buy opinions, we collect them — and measure part of it ourselves.

The tier-list voices (factor "opinion")

SourceTypeNotable
game8.ggMajor guideLeads the board, also provides builds, teams, engines, discs
prydwen.ggCurated guideMaintained monthly, also rates Defense/Anomaly roles
icy-veins.comGuide collectionEditorial tier list
u7buy.comEditorial blogIndependent second opinion
mobalytics.ggTheorycrafting platformIdiosyncratic, data-driven voice

Every source is fetched fresh on every refresh (6-hour cycle + manual button). When sources drop out (Cloudflare, redesign), we keep computing with the remaining ones — confidence (see 4.2) then automatically decreases.

The measurement sources (factors "damage", "stun", "team")

SourceWhat we measure from it
nanoka.cc (datamine)Exact skill multipliers (Ultimate, EX, Basic, Daze) per agent — the basis for damage and stun
game8 build pages (57 agents)All recommended teams per agent — the basis for team fit
Official news (HoYoverse, HoYoLAB, Gematsu)Version, banners, events — controls which agents may be ranked at all (beta agents without gameplay data get no verdict)


2. From tier list to number: the consensus

2.1 Tier → points

Each voice slots agents into grades (T0 through T4). We translate every grade into points:

T0 = 95    T0.5 = 88   T1+ = 82   T1 = 75   T1.5 = 68
T2 = 60    T2.5 = 52   T3 = 42    T4 = 30

2.2 Median instead of mean

For every agent we collect the point values of all sources that rank it and take the median (the middle value), not the average.

Why median: a single outlier voice (e.g. u7buy says T0.5 while everyone else says T2) must not pull a rating up. The median ignores outliers as long as the majority agrees.

2.3 The consensus rank (matters for the verdicts)

Besides the median point value we compute for every agent the consensus rank: the position it would get if we sorted purely by the guide median — i.e. "where would the guides place this agent?". On point ties (many agents share the same median value) each gets the average rank of the group, so that sort order never arbitrarily decides verdicts.

The consensus rank is pure context info ("where would the guides put them?"). The yardstick for UNDERRATED/OVERRATED since the engine refactor is delta_own (see 5) — the distance of our measurements to the role midpoint, not the rank distance.


3. Our own measurements

This is the part no guide can simply copy — here the page earns its own opinion.

3.1 Damage (dpr)

From the nanoka datamine we read the exact skill multipliers per agent:

raw = Basic×0.25 + EX×0.40 + Chain×0.25 + Ult×0.10×decibel factor + Assist×0.10

Rotation model with decibel reality (the team shares 3000 DB): the Ultimate only counts at 10% — and full decibels only for Attack and Anomaly carries (they ult themselves), all other roles only ×0.2. Every category is captured as a chain sum: all hits of the combo, but per hit slot only the strongest variant (uncharged/charged/level-1/2/3 and parry variants are alternatives — exactly one runs in combat). Multiple basic chains (Nicole, Soldier 11) = combo variants → the strongest counts. Stance skills (Yanagi) count as basics. Defensive assist: highest parry tier (Light/Heavy/Chain) × 0.5 — reactive, situational. The raw value is then converted to a percentile within the role group (see 4.1): 90 means "deals more damage than 90% of the agents in its role".

3.2 Stun (daze)

The same datamine provides the daze multipliers. Again: raw value → percentile within the role group. For stun agents this is the core value; for other roles it barely matters.

3.3 Team fit (syn)

We count in how many of the best teams an agent appears — but weighted:

1. Every team on a build page counts with the strength of the page owner (its median tier). A team on Miyabi's page is worth more than the same team on a T3 page. 2. Budget/F2P comps only count 25% — when an "Anby Budget" comp names Anby, that is placeholder advice, not a meta statement. 3. Bayes shrinkage: an agent appearing in only 2 teams so far (new agents!) is pulled cautiously toward the mean instead of punished: syn = (raw × teams + 50 × 6) / (teams + 6). 4. No cap: earlier the guide consensus used to cap our own team measurement here (max 55 at a T2.5 rating) — a circle that prevented UNDERRATED detection. Removed; budget teams still count only as a strength signal.


4. The rating: four factors, role weights, honesty

4.1 Role groups

Anomaly agents (Burn, Shock, DoT) get their own group, because the damage formula cannot measure their damage fairly: a large part of their damage runs through time-based effects that are not contained in hit multipliers. That is why opinion and team count clearly more for them.

4.2 Confidence (the 0-100% display on the cards)

Not to be confused with strength! Confidence says: how sure the system is about THIS agent.

≥2 voices:  conf = 1 − (mean deviation of all voices from the median / 45)
1 voice:    conf = 0.60 + 0.05 × (number of available voices − 1), capped 0.60–0.85
0 voices:   conf = 0.70 (pure own assessment)
Penalty:    agent in ≤2 teams → conf −0.05

If all 6 sources agree → ~100%. If two sources pull in opposite directions → visibly yellow/red. Confidence directly damps the opinion factor in the score formula — divided sources must not fully carry a rating.

4.3 The role weights (fixed, from role logic)

These weights are deliberately fixed — they are our theory of what a role must deliver. They are NOT automatically fitted to the guides (that would be circular: opinion would confirm itself).

RoleOpinionDamageStunTeam
Attack (DPS)35%30%5%30%
Stun35%10%35%20%
Support30%5%10%55%
Anomaly40%10%5%45%
Defense30%5%20%45%

How to read it: a supporter is defined by team fit (55%), a DPS by damage and opinion, a stunner by stun. A Pearson correlation between our measurements and the guide consensus runs as pure diagnostics (see META LAB → CALIBRATION) — it shows where our numbers and the guides diverge, but controls nothing.

4.4 The score formula

score_raw = Σ( factor value × weight ) / Σ( weights )
score     = score_raw × conf + 50.0 × (1 − conf)

Confidence no longer damps the opinion factor in the denominator (that let the relative weight of raw data RISE on source disagreement — mathematically backwards). Instead: global Bayes shrinkage toward the neutral middle 50 — everything uncertain = cautious statement. If a factor is completely missing (e.g. no voice), its weight is automatically redistributed to the rest — no hole appears.

4.5 Smoothing: EMA history

Daily fluctuations (one source changes a grade today) must not throw the list around:

today's score = 30% yesterday's score + 70% fresh computation

Trend arrows (▲/▼) appear from ±3 positions of movement. On a patch change (new version) the history resets completely once — new meta, new merits.

4.6 Positions and bands

Sorted by score; ties share the rank (1, 2, 3, 3, 5...). The DB bands:

S = top 8% (min 3) · A = top 30% · B = top 55% · C = top 80% · T = rest


5. The verdicts

Every verdict compares our measurements with the midpoint of its role (delta_own):

VerdictConditionMeaning
ELITE (checked first)score ≥78 AND top 12%Top with us AND with the experts — the safe pick
UNDERRATEDdelta_own ≥ +8.0 and score ≥50Better than its reputation — our measurements argue against the guides
OVERRATEDdelta_own ≤ −8.0Weaker than its reputation — the guides overlook something
NEWNo source has ranked itToo new — assessment comes from our numbers alone
STANDARDEverything elseWe largely agree with the experts

delta_own = by how much our own measurements (damage, stun, team — role-weighted) raise/lower the agent versus the midpoint of its role (50 = role average of the normalized values). Scale-free — no longer a rank delta that produced false alarms in the mid-field (20 agents within 3 points) and mathematically never allowed UNDERRATED at the top (max 4 positions apart). No tier grouping: even a lone agent in its tier gets a true distance to the role midpoint.

Example: Nicole, supporter. Her measurements (team fit 92/100 — she appears in 33 of the best teams) lift her clearly above the supporter midpoint: delta_own ≈ +12 → UNDERRATED. No rank comparison, no tier view — just the question: "how far above the role average do her measured values lie?"

The ±8-point threshold exists because smaller distances are noise — only a real point distance between our measurements and the role average is a disagreement that deserves a name.

POWER-100 (v4 scoring, 2026-08-27)

The composite is now a weighted geometric mean, fully transparent:

score = Π(factor^weight) (factors 5..100; missing factor = neutral 50)

Properties: all factors 50 → 50 (average), all 95 → 95 (peak); a weak factor drags more than in an arithmetic mean — a kit with a hole loses combat power, which is exactly the realism we want. No Bayes shrink toward the middle, no EMA smoothing, no clamp saturation: the number is 100% reproducible from the displayed factors and weights.

Class profiles are data-driven, not hand-picked. Every specialty (including future ones like Rupture, which arrived with 3.1) forms its own measurement peer group. The damage/stun weighting per class is derived from the class's median kit orientation, measured in GLOBAL percentiles: where does this class sit across all agents in damage output vs daze output? Stun classes legitimately weigh daze higher, attack classes damage — the game data itself defines the profile:

Classopiniondamagestunteam
Attack0.300.2240.1960.28
Stun0.300.1570.2630.28
Support0.300.1790.2410.28
Anomaly0.300.2400.1800.28
Defense0.300.2250.1950.28
Rupture0.300.1660.2540.28

(These numbers regenerate on every refresh — a new class automatically gets its own profile the moment ≥1 agent of it exists.)

Hollow Index vs pull economics: the index measures STRENGTH in a vacuum — pull decisions are investment economics (role, team cost, engine, opportunity cost). A high score is never a pull recommendation by itself; the SIGNALS view weighs both.


6. Why you should trust us — and where not

In favor

1. Transparency down to the formula. This document, the per-role weights (META LAB → CALIBRATION), the factor bars and the confidence bar on every card — all visible, nothing behind a blackbox. 2. Five independent voices instead of one. No single editor decides the opinion. Where the sources fight, you see it (confidence drops, the divergence table lists every deviation). 3. Own measurement instead of copy. Damage and stun come from exact datamine numbers, team fit from real build recommendations — not from retyping a tier list. 4. Admitted limits. Every WHY-box names the weaknesses of its own rating (measurable limits below). A system that knows its limits does not lie. 5. No fashion swings. EMA smoothing + patch reset: the list reacts to real meta changes, not to a source's daily mood. 6. Disagreement is allowed. UNDERRATED/OVERRATED openly show where we contradict the guides — with a per-agent justification. A page that only mirrors them would not need us.

Limits (deliberately documented)

  • Damage measures burst, not everything: time-based effects (Burn, Shock, DoT) and
  • follow-up mechanics are not contained in hit multipliers — anomaly agents therefore have their own weight profile, but the damage value remains a lower bound for them.

  • Rotation and field time do not enter: an agent who plays slower but hits harder looks
  • better in the damage value than it is in real runs.

  • Defenders have their own role (opinion .30 / damage .05 / stun .20 / team .45) — but
  • shields/survival stay invisible, there is no shield data in the datamine.

  • New agents (few teams, possibly no voice) have damped confidence and rough team
  • estimates — the WHY box says so explicitly.

  • Beta agents (not unlocked yet, no build data): no verdict, just NEW.
  • The sources themselves can be wrong. The consensus is the starting opinion, not the
  • truth — exactly why we measure alongside.


    7. The pipeline in one picture

      6 tier lists ───► tier→points ───► MEDIAN ───► consensus rank (exp_rank)  ┐
                                                                                │
      nanoka datamine ─► skill multipliers ─► dpr/daze percentiles ─────────────┤
                                                                                ├─► role weights
      game8 builds ───► teams (tier-weighted, ─► syn (Bayes-smoothed) ──────────┤     (fixed)
                        budget teams count 25%)                                  │
                                                                                └─► SCORE
      confidence (voice agreement, team coverage) ──► Bayes shrink →50           │
      EMA (30% yesterday + 70% today) ──► smoothed score ◄────────────────────────┘
                                                                                  │
      delta_own (measurements vs role midpoint) ──► UNDERRATED / OVERRATED /
                                                   ELITE / STANDARD / NEW
    


    8. Verification & reproducibility

  • Every refresh stores HTML snapshots of all sources (snapshots/<date>/) — every number
  • is traceable back to the source.

  • The complete code (scripts/refresh.py, ~1600 lines, commented) lives in the repo.
  • Git history: every methodology change is committed and diffable.
  • Every 6 hours the refresh runs automatically (plus watchdog); the page shows the
  • data stamp everywhere.