🆔 Coding

UUIDs and Hashes: Two Things That Look Random and Are Not Interchangeable

A UUID is an identifier. A hash is a fingerprint. Both are hex strings of fixed length, and using one where the other belongs causes distinct, predictable problems.

Both produce a fixed-length string of hex that looks like noise. That is where the similarity ends. A UUID is generated to be unique and carries no information about anything else. A hash is derived from input and is a deterministic fingerprint of it — the same input always gives the same hash, which is the entire point and also the entire problem in one specific case.

Here is what each is for, which variants matter, and the two mistakes that keep recurring.

What a UUID actually guarantees

A UUID is 128 bits, written as 32 hex digits in five hyphenated groups: 8-4-4-4-12. The format is fixed; what varies is how the bits are chosen, which is the version.

Two positions are reserved. The first digit of the third group is the version number. The first digit of the fourth group encodes the variant and is always 8, 9, a or b for RFC-4122 UUIDs. So in:

f81d4fae-7dec-41d0-a765-00a0c91e6bf6

the 4 says version 4 and the a says RFC 4122. Those two characters are the cheapest validity check you can run, and they are why a random 32-hex-digit string is not a valid UUID.

Version 4 is 122 random bits (128 minus the 6 fixed ones). The uniqueness guarantee is probabilistic, not absolute — but the probability is not a reason for concern. Generating a billion v4 UUIDs a second for a century gives roughly a 50% chance of one collision. You will hit other limits first.

v1, v4 and v7: why the choice matters for databases

Version 1 embeds a timestamp and a node identifier, usually derived from a MAC address. The timestamp is a 60-bit count of 100-nanosecond intervals since 15 October 1582 — the date the Gregorian calendar was adopted, which is the kind of detail that makes UUID v1 memorable. It is sortable by creation time, and it leaks the machine that generated it, which is why it fell out of favour.

Version 4 is pure randomness. No information leaks, and no two generators need to coordinate. Its weakness is physical: as a primary key in a B-tree index, every insert lands at a random point in the tree. That fragments pages, inflates the index and destroys the locality that makes sequential integer keys fast. On a large table this is a measurable write penalty.

Version 7, standardised in RFC 9562 in 2024, fixes exactly that: a 48-bit Unix millisecond timestamp followed by random bits. It is sortable, roughly sequential, and leaks only the creation time — which for most records is already a column. If you are choosing a UUID version for a database key today, v7 is the answer; v4 remains right for tokens, correlation IDs and anything where ordering would be a leak.

Versions 3 and 5 are different in kind: they hash a namespace and a name to produce a deterministic UUID. The same name in the same namespace always gives the same UUID, which is useful for deriving stable identifiers from existing keys. v5 uses SHA-1, v3 uses MD5; prefer v5.

What a hash is for

A cryptographic hash maps input of any size to a fixed-size digest, with three properties that matter:

  • Deterministic — the same input always gives the same digest.
  • One-way — given a digest, you cannot feasibly recover the input.
  • Collision-resistant — you cannot feasibly find two inputs with the same digest.

Changing one bit of input changes roughly half the output bits. That avalanche property is what makes a hash useful as a fingerprint: report.pdf and report.pdf with one comma moved have digests with nothing visibly in common.

The algorithms you will meet:

  • MD5 (128-bit) — collisions are trivially constructible. Fine for a non-adversarial checksum, unacceptable for anything security-relevant.
  • SHA-1 (160-bit) — broken in practice since 2017. Still appears in legacy protocols and Git object IDs.
  • SHA-256 (256-bit) — the current general-purpose default.
  • SHA-512 (512-bit) — same family, larger digest, often faster on 64-bit hardware.
  • CRC-32 — not cryptographic at all. A 32-bit error-detection code, designed to catch transmission corruption, trivially forgeable. Useful, but never for integrity against an adversary.

Why SHA-256 is wrong for passwords

This is the mistake worth the most words, because the reasoning is counterintuitive: SHA-256 is unsuitable for password storage because it is fast.

Modern hardware computes billions of SHA-256 digests per second. An attacker holding a stolen table of SHA-256 password hashes does not try to reverse them — they hash candidate passwords until they match. At those speeds, every password in any wordlist falls immediately, and most human-chosen passwords are in a wordlist.

Password hashing needs a function that is deliberately expensive and tunable, so defenders can set a cost that keeps verification imperceptible while making bulk guessing impractical:

  • Argon2id — the current recommendation. Tunable in time, memory and parallelism; the memory cost is what defeats GPU and ASIC attacks.
  • scrypt — also memory-hard, widely available.
  • bcrypt — older, well understood, still acceptable. Note its 72-byte input limit.
  • PBKDF2 — the weakest of the four, as it is not memory-hard, but it is in every standard library and FIPS-approved.

All of them salt by default. A salt is unique random data per password, stored alongside the hash; it stops one precomputed table from attacking every account at once, and stops two users with the same password from having the same hash. A plain SHA-256 has no salt and no cost parameter, so it provides neither defence.

The short version: use SHA-256 to fingerprint a file, and Argon2id to store a password. They are not substitutes in either direction.

The encoding detail that breaks checksums

A hash is computed over bytes, not characters. So "the SHA-256 of this text" is underspecified until you say how the text becomes bytes.

For pure ASCII it does not matter. The moment a non-ASCII character appears, UTF-8 and UTF-16 produce different byte sequences and therefore different digests. The word café is 5 bytes in UTF-8 and 8 in UTF-16LE; the SHA-256 values have nothing in common.

This is why two tools can disagree about the hash of the same visible string, and the disagreement always starts with a non-ASCII character. UTF-8 is the de facto standard — if a digest does not match, check whether one side is using UTF-16, and whether a trailing newline has crept in. A file that ends in a newline hashes differently from one that does not, and text editors differ on whether to add one.

Generating and checking locally

The UUID generator produces v4 UUIDs from the browser's cryptographic random source — not Math.random(), which is not suitable for identifiers that must not be guessable — and v1 UUIDs with a correctly encoded timestamp, in bulk and in the format you need.

The hash generator computes SHA-1, SHA-256, SHA-512 and CRC-32 over the UTF-8 bytes of your input, via the Web Crypto API, and will verify a digest you paste in. It runs in the browser, which matters if the thing you are hashing is not something you would upload.

For passwords, neither of these is the right tool and neither pretends to be — use the password generator to create one, and a real password-hashing function on your server to store it.

Frequently asked questions

Can two UUIDs collide?

In principle, yes; in practice, no. A version 4 UUID has 122 random bits. Generating a billion a second for a century gives roughly a 50% chance of a single collision — you will encounter other constraints long before that one.

Should I use UUID v4 or v7 as a database primary key?

v7. It begins with a millisecond timestamp, so inserts land at the end of the index instead of scattering through it, which avoids the page fragmentation and write amplification that v4 causes in a B-tree. Keep v4 for tokens and anything where creation order would leak information.

Why is SHA-256 bad for storing passwords?

Because it is fast. Hardware computes billions of digests per second, so an attacker with a stolen hash table simply guesses. Use Argon2id, scrypt or bcrypt — functions deliberately expensive and memory-hard, with a tunable cost — and always with a per-password salt.

What does a salt do?

It is unique random data stored with each password hash. It stops a single precomputed table from attacking every account at once, and ensures two users with the same password do not share a hash. Plain SHA-256 has no salt, which is part of why it is unsuitable.

Why do two tools give different hashes for the same text?

A hash is computed over bytes, and text becomes bytes via an encoding. UTF-8 and UTF-16 give different digests for anything non-ASCII. Check the encoding first, then check for a trailing newline — a file ending in one hashes differently from one that does not.

Everything on ToolYard runs in your browser. No uploads, no accounts, no limits.

Browse all tools →