
What every number here actually means
This tool makes a lot of specific claims. This page explains each one, what it's based on, and — just as importantly — where the limits are. If something here can't be traced to a real rule or weight, it says so.
What this is (and what it isn't)
In August 2026, X open-sourced the code that decides what appears in the For You timeline. This tool is built on that code: the ranking weights, the filters, and the visibility rules are read directly out of this repository rather than guessed at from social-media folklore.
The honest limitation, up front
X's real ranking runs Phoenix, a learned transformer that reads a specific viewer's engagement history to predict how likely that person is to reply to, share, or mute your post. We don't have that model, your account's history, or a viewer. So this tool estimates those probabilities from surface features of your draft — length, hooks, links, formatting — and then blends them using the real production weights. The weights and the rules are accurate. The probabilities are a transparent stand-in. Treat the score as directional coaching, not a promise of reach.
That's also exactly why the Track record feature exists — it checks the tool's predictions against your real posts so you don't have to take any of this on faith.
The BESC Score
One number, 0–100. It's built the same way X builds its own ranking score, then squashed into a readable range.
- 1Estimate a probability for each action a viewer might take — reply, like, repost, share, follow you, or mute/block/report you.
- 2Multiply each by its real production weight and sum them. This is X's actual formula:
Final Score = Σ (weight × P(action)). - 3Apply the author-diversity multiplier if you've posted recently (see below).
- 4Squash to 0–100 so it's readable. A weighted sum is unbounded; the score is not.
What the grades mean
| Score | Grade | Roughly |
|---|---|---|
| 82–100 | Excellent | Strong hook, a real reason to reply, nothing risky |
| 58–81 | Strong | Solid post, usually one lever left unpulled |
| 38–57 | Decent | Publishable, but leaving reach on the table |
| 20–37 | Weak | Little reason for anyone to act on it |
| 0–19 | High Risk | Likely tripping a spam/negative signal — check the risk panel |
The two multipliers shown under the gauge
- Author diversity ×
- Each additional post from you in the same window is multiplied by
(1 − 0.25) × 0.5^k + 0.25. Your 2nd post scores at ~62%, and it floors at 25% by the 4th. Spacing posts out is worth more than most people think. - Out-of-network ×
- Your own followers see the post scored at full strength. For everyone else it's multiplied by
0.75— and on a topic-matched recommendation surface it's0.5. That surface also runs a much longer list of drop rules your followers never hit.
Every input explained
These change the score because they change how the real algorithm treats the post. None of them are cosmetic.
- Media type (Text / Photo / Video / GIF)
- Media unlocks scored actions text can never earn — photo-expand, video-open, and video-quality-view. One catch worth knowing: a video under 10 seconds has its video-quality-view weight forced to exactly 0, no matter how good it is.
- This is a reply
- Structurally the most expensive toggle here. Replies and reposts are removed entirely from recommendations to anyone who doesn't already follow you — not downranked, excluded. Even shown to your own followers they're rescored at the same 0.75× as out-of-network content. Replies are also excluded from the cold-start boost and the mutual-follow bonus below.
- Mostly mutual-follow audience
- If people who follow you also get followed back, an original post gets
+15.0added to its reply weight — the single largest situational boost in the model, taking reply from 5.0 to 20.0. It never applies to a reply, which is why this is greyed out when "This is a reply" is checked. - Contains sensitive media
- Costs more than the blur suggests. Followers see it behind a click-through interstitial — but for everyone else it's dropped from recommendations entirely. It can also accumulate into an account-level label if roughly 3 of your last 5 posts are flagged.
- Verified / Premium checkmark
- Raises the character limit from 280 to 4,000 and changes which length-related tips apply. It is not treated as a ranking boost, because the open-sourced ranking code doesn't contain one.
- Posts already sent in this window
- Drives the author-diversity multiplier described above.
- @handle + follower count
- Used for the cold-start check, and to remember your track record. Under 1,000 followers, a post under 24 hours old and under 1,000 views can be lifted in ranking — but only if it's already ranking in the top 85% of a viewer's candidates on its own merits, and it's lifted toward roughly rank 15, not to the top of the feed.
Signal breakdown
The full ledger behind the score: every action, its real weight, the estimated probability, and what it contributed. These weights are read straight from home-mixer/params/param.rs.
| Action | Weight | Worth knowing |
|---|---|---|
| Share via copy link | +20.0 | The highest single weight in the entire model |
| Reply | +5.0 → +20.0 | With the mutual-follow boost on an original post |
| Quote post | +5.0 | |
| Share via DM | +5.0 | |
| Follow you | +4.0 | |
| Share (generic) | +2.0 | |
| Repost | +1.0 | |
| Like | +0.5 | The lowest positive weight — chasing likes is chasing the least valuable action |
| Not interested | −43.2 | |
| Block | −31.2 | |
| Mute | −58.8 | |
| Report | −234.0 | The largest weight in the model, by far, and it's negative |
The asymmetry is the whole point
A report is worth 468 likes in the negative direction. Anything that nudges people toward mute, block, or "not interested" — shouting, spammy formatting, engagement bait — costs far more than the engagement it might buy.
Visibility-filtering risk
Ranking decides order. Visibility filtering decides whether your post can be shown at all — a separate system with three possible answers: allow, show behind an interstitial, or drop. Each flag here cites the specific rule file it comes from.
- Link / URL verdict
- Shorteners, raw IP links and certain cheap TLDs draw scrutiny. A "low quality" verdict gets downranked; an "unsafe" one is a hard drop for every non-follower. It isn't contained to one post either — when a domain's reputation flips, past posts sharing it get relabelled too.
- Templated / copy-paste phrasing
- "Follow for follow", "link in bio", "RT if you agree" — the exact fingerprint duplicate-text spam detection looks for across accounts.
- ALL-CAPS / !!! bursts / hashtag stuffing
- Pushes up the three most negative weights in the model. Note that a word you also use as a hashtag (a ticker or brand name) is treated as a name, not as shouting.
- Post age
- Anything older than 48 hours stops being eligible for For You ranking altogether — a hard exclusion before scoring even runs.
- Sensitive media & repeat NSFW
- Interstitial for followers, dropped from recommendations for everyone else.
Checks that pass are collapsed under "other checks passed clean" so the panel shows you what actually needs attention.
Tips
Ranked by impact, and each one is tied to a weight rather than to taste. The highest-impact tips are almost always the same two: give people a concrete reason to reply (worth 10–40× a like), and give them something specific enough to be worth copy-link sharing (the single heaviest action).
Not all reply hooks are equal
A question tied to the specifics of your post scores substantially higher here than a bolt-on closer like "What do you think?". That's deliberate, and it was a bug worth fixing: the scorer used to reward the generic closer more, which quietly pushed every optimized post toward the same lazy ending. Two reasons it runs the other way now — a question that asks nothing gives a reader no reason to answer, and thousands of posts ending in an identical tail is precisely the templated-text pattern duplicate-text spam detection looks for across accounts.
Others cover writing craft — filler words, passive voice, weak openers, stock AI phrasing. Those are flagged as general craft signals, not as repo-cited weights, and the tool labels them that way rather than dressing them up as algorithm rules.
Optimize for the algorithm
A deterministic, meaning-preserving pass. It applies mechanical fixes and keeps each one only if the score measurably improves, verified with the same scorer used everywhere else. "Optimized" here provably means higher-scoring, never just "reworded".
| Fix | Why |
|---|---|
| Toned down !!! / ??? bursts | Feeds the same spam penalty as ALL-CAPS |
| Fixed ALL-CAPS shouting | Drives mute / report / not-interested propensity |
| Trimmed hashtags to 2 | Beyond a couple reads as stuffing and dilutes the post |
| Removed templated CTA phrasing | The fingerprint duplicate-text spam detection catches |
| Cut filler/hedge words | Dilutes a claim without adding information |
| Added a reply hook | Reply is worth 10–40× a like. A deterministic pass can't write a question about your specific post, so this is a varied fallback — a specific question scores materially higher |
| Trimmed to the character limit | A hard constraint — applied whether or not it raises the score |
It won't rewrite a brand name or ticker that you also use as a hashtag, and it won't invent facts, because it only ever deletes or reshapes text you already wrote.
AI rewrite & generate from an idea
Two optional layers on top of the deterministic optimizer. Both are available only if the deployment has an AI provider configured, and every AI output is still scored and gated by the same deterministic scorer — the AI only ever suggests, it never decides.
The model isn't just told "make this engaging". It's given a working brief on the ranking system, assembled from the same constants the scorer uses: the full weight table and what the ratios imply, the structural limits that cap reach before wording matters (reply/repost exclusion, the out-of-network discount, author-diversity decay, the 48-hour cutoff), the label risks that can drop a post outright, and the constraints of your specific post — what media is attached, whether it's a reply, how many times you've posted in this window. That's what lets it write something worth copy-link sharing rather than tacking a question onto the end.
- AI rewrite
- Rewrites your existing draft, preserving your facts and voice. Candidates are only shown if they score higher than the mechanical result. If one beats it by 5+ points it's applied automatically, with a one-click Undo and the fact-check warning kept in view.
- Generate from an idea
- For when you have something to say but no draft. Give it rough notes and it writes complete posts, each already run through the deterministic optimizer and scored, best first.
Read AI output before you post it
An AI can rephrase a claim in a way that changes its meaning — turning a pending status into a finished one, for example — even when explicitly told not to. Generated posts are the higher risk of the two: there's no original text to check them against. The prompts forbid inventing facts, names, numbers and links, but a higher score only means it fits the algorithm's signals better, not that it's true. Verify every name and number yourself.
Import a live post
Paste any x.com/…/status/… link to score a post that already exists. It pulls the real text, media type, author follower count, verified status, post age and real engagement numbers, and fills in the toggles it can determine for itself — including whether the post was a reply and whether it was marked sensitive.
Useful for scoring a competitor's post, or for going back over your own to see what the tool would have said before you published it.
Track record
The part that keeps the rest of this honest. Everything above is a prediction. This checks the predictions against reality.
- 1Hit Track on a draft. It's saved with the score you're looking at.
- 2Publish the post on X, in your own composer, as normal.
- 3Hit Check for results. Your recent timeline is matched against saved drafts — tolerant of last-minute wording tweaks and X's link rewriting — and the real views and engagement are pulled once the post has had a couple of hours to accumulate them.
What it will and won't tell you
Once you have 6 measured posts, it splits them at the median predicted score and compares the real numbers of the higher-scoring half against the lower-scoring half. If those two columns look the same, the score isn't predicting anything for your account — and the panel will say exactly that rather than spinning it.
Why it refuses to answer early
Engagement data is noisy enough that a "pattern" drawn from three posts would be invented rather than observed. So nothing is claimed below 6 measured posts, per-fix comparisons need at least 3 posts on each side, and everything uses medians rather than averages so a single post that happens to take off can't manufacture a trend. A tool like this is only worth having if you can trust it when it says "not yet".
A fix is only credited to a post when the tracked text is exactly the optimizer's output — edit after optimizing and nothing is recorded for that post, because a wrong attribution would quietly corrupt the very data this exists to build.
Learning from your history
Learn from my history reads your published posts and their real numbers in one go, rather than waiting for tracked drafts to accumulate. Views and per-action counts are both public on a published post, so replies ÷ views is a directly measured action rate — real ground truth, not an estimate.
With enough of those, the guessed probabilities behind your score get replaced by ones fitted to your own results, using ridge regression over the same features the scorer already extracts. It's the same shape as what Phoenix does — predict a probability per action, then blend with the real production weights — just learned at the level of "posts like this one" rather than per individual viewer.
Why it can't be Phoenix itself
Phoenix takes a viewer, that viewer's engagement history, and a candidate post, and predicts what that specific person will do. An unpublished draft has no viewer, and no API exposes the private engagement history of everyone who might see your post — so the model is out of reach by construction, not for lack of access. No trained weights are published either; what ships is training code and synthetic data generators. The aggregate question this tool asks — "how will this post do?" — is a different and more tractable one.
Three guards keep it honest. Nothing is fitted below 40 posts. Every fitted action must clear cross-validated R² before it's used at all, so a model that merely memorised your history is rejected rather than shipped. And the fit is blended with the original heuristics in proportion to how much data backs it — 60 posts nudges the score, 500 largely drives it.
What's learned is relative: "this post should do about 1.8× your typical reply rate", applied to the existing scale. Real reply rates are ~1% of views while the priors sit near 0.4 — they were tuned to spread the 0–100 range usefully, not to be literal probabilities. Substituting real magnitudes would drop everyone's score by a dozen points and tell them nothing, so calibration reorders drafts rather than rebasing the scale.
Tracking needs a database to be attached to the deployment. Without one it reports itself as unavailable and everything else works exactly as before.
Where the numbers come from
Every weight, threshold and rule cited in this tool is read out of the open-sourced algorithm in this repository. The main ones:
| What | Source |
|---|---|
| Action weights, OON discount, cold-start params | home-mixer/params/param.rs |
| Score arithmetic, author diversity, reply/repost rescoring | home-mixer/scorers/ranking_scorer.rs |
| Cold-start boost eligibility and targeting | home-mixer/scorers/author_cold_start.rs |
| 48-hour age cutoff, reply/repost OON exclusion | home-mixer/filters/ |
| Drop / interstitial rules and their evaluation order | visibility-filtering/rules/registry.rs |
| Label definitions and their real effects | under-the-hood/strato/lib/underTheHoodLabels.strato |
| Spam / URL-verdict labelling rules | botmaker-rules/scarecrow/bot/ |
The scoring implementation is fully commented with a citation for every number in besc-engagement-checker/lib/scoring.ts. If you think one of them is wrong, it's all readable — and worth telling us about.