Deduplication in PHP: what it means and how to do it
| What deduping means | |
| Same by what? Equality vs identity | |
| Deduping arrays | |
| Deduping file loads and definitions | |
| Deduping at the database | |
| Checklist |
What deduping means
Deduplication is removing repeated items so each distinct thing is
processed, stored, or shown exactly once. Every dedup decision answers
two questions:
1. What counts as the same? Byte-identical rows? Same e-mail
with different case? Same helper file reached through two paths?
The answer defines the comparison key.
2. Where do you enforce it? At input (reject early), at storage
(constraints), or at output (filter when reading). Enforcing at the
wrong layer is the classic source of duplicates: the code looks
correct, the data is not.
Same by what? Equality vs identity
PHP compares loosely with == and strictly with === (value plus type). Dedup keys should almost always be strict: "10" == 10 is true, and a loose key merges two things the user sees as different. Normalize first (trim, case-fold, cast), then compare strictly — the normalization is your definition of sameness.
Deduping arrays
array_unique() keeps the first occurrence of each value — but by default it compares values as strings, so 10 and "10" collapse into one. Pass a sort flag (SORT_REGULAR, SORT_NUMERIC) to match your definition of sameness, and note it preserves original keys — reindex with array_values() if the gaps matter.
$ids = array_values(array_unique($ids, SORT_REGULAR));
For large inputs the idiomatic O(1) pattern is keys-as-set: assign each item as an array key and the engine dedupes for you, preserving insertion order on read-back:
$seen = [];
foreach ($rows as $row) { $seen[$row['email']] = $row; }
$unique = array_values($seen); // last occurrence wins
Deduping file loads and definitions
require_once / include_once dedupe by resolved file path — two spellings of one file load once, but two different files defining the same function both load, and the second definition fatals. Paths dedupe files; only function_exists() / class_exists() guards dedupe definitions:
if (!function_exists('send_report')) {
function send_report() { ... }
}
Full story with experiments: require vs include (path rules) and Cannot redeclare get_urls() (one name, three engine trees, 249 fatals).
Deduping at the database
Application-level dedup races: two requests can both check, both find nothing, both insert. The database is the only layer that sees every writer, so uniqueness that matters belongs there — SELECT DISTINCT for reading, UNIQUE constraints for storing, and explicit conflict policy on write (INSERT ... ON CONFLICT DO NOTHING in PostgreSQL, INSERT IGNORE / ON DUPLICATE KEY UPDATE in MySQL). PHP-side filtering stays useful for display, never as the guarantee.
Checklist
1. Define sameness first: normalize, then compare strictly.
2. Arrays: array_unique() with an explicit
flag, or keys-as-set for large inputs.
3. Files: _once dedupes paths;
function_exists / class_exists dedupes
definitions.
4. Data that must be unique lives behind a database constraint, not
an application check.
5. After renames and deploys, purge OPcache — stale bytecode
replays yesterday duplicates.
Article author: Arthur Isaev