Ask AI about me

Choose an assistant to ask about my work.

Opens an external service. AI answers can be inaccurate.

←Back to Blog
July 22, 20269 min readsecurityprivacydevelopmentpasswordsprojects

Password Strength Meters Are Lying to You

A green bar often means your password satisfied a checklist, not that an attacker would struggle to guess it. My own checker proves how misleading that can be.

Type Password1! into a typical password-strength meter and watch what happens.

The bar moves to green. The label says Strong. The form is satisfied because the password has an uppercase letter, lowercase letters, a number, a symbol, and enough characters.

An attacker is also satisfied because Password1! is exactly the kind of predictable transformation a password cracker tries early.

The meter did not measure strength. It graded obedience.

I know because I built one of these meters. My Password Strength Checker is a tiny client-side project that evaluates passwords in real time. It is fast, private, and visually clear. It also uses the same character-class logic I now think developers should stop presenting as security.

The project is useful. The score is not truth.

What my checker measures

The implementation is only 76 lines of JavaScript, which makes its assumptions easy to see.

It awards points when a password:

  • Contains at least eight characters
  • Contains at least twelve characters
  • Includes a lowercase letter
  • Includes an uppercase letter
  • Includes a number
  • Includes a symbol

It subtracts one point for three repeated characters, clamps the result from zero to six, and converts that score into labels from “Very Weak” to “Very Strong.” Everything happens locally in the browser. Nothing is stored or transmitted.

That last property is genuinely good. A password meter should not need to send a secret to a server just to color a bar. The problem is what the local code claims to know.

Under this system, Password1! receives five of six points and becomes Strong. A long phrase such as correct horse battery staple receives points for length, lowercase letters, and the spaces counted as symbols, but loses points for having no uppercase letters or numbers. A randomly generated 20-character lowercase password receives only three points even though, if it was truly chosen uniformly, its guessing space would be enormous.

The meter rewards visual variety. Attackers care about predictability.

Those are not the same property.

Character classes are security theater

Uppercase-letter-number-symbol rules feel mathematical because they appear to expand the set of possible characters. If every position in a 12-character password were chosen independently and uniformly from a large alphabet, that expanded alphabet would matter.

Humans do not choose passwords that way.

When a site demands one uppercase letter, people often capitalize the first letter. When it demands a number, they add 1 or a year at the end. When it demands a symbol, they append !. The theoretical search space grows while the practical set of human choices stays concentrated around predictable patterns.

Attackers do not begin with every possible string in alphabetical order. They begin with leaked passwords, dictionaries, names, dates, keyboard paths, common phrases, and mutation rules. They already know that password becomes Password1!. A composition checklist congratulates the user for making the exact substitutions the attacker expected.

Current NIST password guidance explicitly says verifiers should not impose character-composition rules. For a password used as a single authentication factor, it requires a minimum of 15 characters, recommends permitting at least 64, and requires checking new passwords against a blocklist of common, expected, or compromised values.

That is a very different philosophy from “add a symbol to make the bar green.”

Entropy is real and usually misused

Password discussions love entropy because it turns strength into a clean number of bits.

For a password generated uniformly at random, the calculation is useful. If each of L positions is independently selected from an alphabet of N possible characters, the theoretical entropy is:

entropy = L × log2(N)

Twenty characters randomly selected from 26 lowercase letters represent about 94 bits. That is why the all-lowercase random password can be much stronger than the shorter password that checks every visual-complexity box.

But the formula describes the generation process, not the appearance of the final string.

I cannot look at Tr0ub4dor&3 and honestly calculate its entropy by pretending every character was sampled independently from 95 printable characters. A human chose a dictionary word, replaced obvious letters, capitalized the beginning, and added predictable characters. The apparent alphabet says almost nothing about the probability of that choice.

NIST's own explanation notes that estimating entropy for human-chosen passwords is difficult. That is why length, blocklists, secure storage, and attack controls are more practical than decorating a meter with a fake bit count.

Entropy is not a vibe. It belongs to a known random process.

The better question is “how many guesses?”

A more honest meter tries to estimate password guessability: approximately how many attempts a capable attacker would need under a particular attack model.

That requires recognizing patterns composition meters ignore:

  • Common passwords and previously breached passwords
  • Dictionary words and familiar phrases
  • Names, dates, and context connected to the user or service
  • Keyboard walks such as qwerty and 1qaz2wsx
  • Repeated and sequential characters
  • Predictable capitalization and substitutions
  • Combinations attackers have already learned from real leaks

Tools such as zxcvbn model these patterns and estimate a guess count rather than awarding one point per character class. Research on password meters goes further: there is no single perfect guesser or single accuracy metric. Different attackers, training data, online limits, and offline cracking conditions change what “strong” means. The USENIX research on password-meter accuracy is blunt about this—there is no universal silver-bullet measurement.

That uncertainty does not make meters useless. It means the interface should stop pretending a six-point score is a physical measurement.

Online and offline attacks are different problems

A password that survives an online login form may fail quickly after a database breach.

Online guessing can be rate-limited. The service can slow repeated failures, detect suspicious behavior, require another factor, or lock an attack path. Under those conditions, even a moderate password can resist the small number of guesses an attacker receives.

Offline cracking begins after an attacker steals password hashes. There is no login server left to enforce rate limits. The attacker can guess as quickly as the storage algorithm and hardware allow. A properly salted, deliberately expensive password hash makes every attempt cost more; weak or fast storage lets the attacker test guesses at enormous scale.

A meter inside the password field cannot see the server's hashing configuration. It does not know whether the account uses multi-factor authentication. It does not know whether the password was reused on a breached site. It cannot tell whether malware is recording the keyboard or whether the user is about to enter the password into a phishing page.

Calling a password “Very Strong” without a threat model is already overpromising.

What my checker gets right

The project has one privacy property I would absolutely keep: evaluation stays in the browser.

The input listener passes the current string directly to a local function. The function returns a score and feedback. The UI updates. There is no analytics request, account, database, or API call. A user can download the files and run the checker offline.

That architecture is better than a sophisticated meter that leaks the password it evaluates.

Local does not automatically mean safe, because the page itself still has to be trusted. A compromised script can read anything typed into its input. But reducing the system to plain HTML, CSS, and JavaScript makes that trust easier to inspect.

The feedback is also understandable. “Use at least 12 characters” is actionable. The mistake is presenting missing uppercase letters or symbols as if they necessarily indicate weakness.

The meter would be more honest if it separated requirements from estimates. A site may have a compatibility rule. That rule can be shown as satisfied without being added to a fictional security score.

What I would change

I would keep the tool local and replace the composition score with a layered explanation:

  1. Prioritize length. Longer user-chosen passwords and passphrases generally require more guesses, without forcing hard-to-remember punctuation rituals.
  2. Check a local blocklist. Reject common, expected, and compromised passwords rather than congratulating predictable mutations.
  3. Estimate guessability. Detect dictionary terms, keyboard patterns, dates, repeats, sequences, and common substitutions.
  4. Explain the pattern. Say “this is a common password with a predictable suffix,” not merely “Weak.”
  5. Do not punish harmless choices. Spaces and lowercase-only random passwords should not lose points for looking insufficiently complicated.
  6. State the limit. The meter estimates resistance to guessing. It does not protect against reuse, phishing, malware, insecure server storage, or account recovery failures.

I would also recommend a password manager more aggressively. The best answer is usually not teaching a person to invent a clever password. It is generating a long unique one for every account and protecting the manager with a strong memorized passphrase and another authentication factor.

Passkeys can remove the shared-secret problem entirely for supported accounts. A password-strength meter should be honest enough to admit when the best password is no password.

The honest verdict

Most password meters are lying because they measure visible compliance and label it security.

My own checker proves the problem cleanly. It is private, quick, readable, and wrong in a very conventional way. Password1! should not become strong because it collected five character-class points. A long randomly generated lowercase string should not become weak because it refused to perform for the checklist.

The right question is not “how many kinds of characters can I see?” It is “how early would this exact choice appear in a realistic attacker's guesses?”

That answer is harder to calculate, depends on the threat model, and will never fit perfectly into a colored bar.

Security tools should communicate that uncertainty instead of painting it green.

Keep reading