Review duplicate records before choosing what to keep
Select exact identity columns, inspect every group member, then explicitly retain the first or last source occurrence. First and last mean file order, not newest timestamp or best quality. Your original file stays unchanged.
Input: UTF-8, 2 MiB, 20,000 data records, 100 columns, 200,000 total cells including header. Selecting a file replaces pasted text even if rejected. Each CSV partition independently allows 2 MiB including quotes, CRLF, header and possible prefixes; both downloads are suppressed if either fails. Raw JSON has a separate 16 MiB serialized envelope cap and canonical rows once; it does not share the CSV byte cap.
Choose identity columns
If any selected field is empty or whitespace-only, the identity is incomplete: the row is excluded from grouping and always retained unchanged. Legitimately blank identities cannot be deduplicated here. Nonblank spaces, case and leading zeroes remain exact.
Ready for an export.
Duplicate review
All member rows in duplicate groups
Table previews at most 100 member rows. The exact JSON contains every canonical source row and all group references.
Prefix mode adds an apostrophe to formula/control-leading cells, including = + - @ after whitespace or a leading tab/newline. It changes values and is not universal spreadsheet safety. Literal mode retains strings but spreadsheet software may interpret formulas. JSON and review remain raw.
Worked example: matching identity, different notes
Load the six-record example, read and select department and code. Review two groups: source records 2/4 and 5/6, with all four member rows visible; two rows are singletons. Six records reconcile as 0 incomplete + 4 group members + 2 singleton rows. There are four complete identities: 2 groups + 2 singleton identities.
Keep first retains records 2,3,5,7 and removes 4,6. Keep last retains 3,4,6,7 and removes 2,5. Both partitions preserve source order and full original headers. The tuples (ab,c) and (a,bc) differ, as do North and north. Matching selected columns does not make the differing note cells identical.
Limits and data handling
No default winner is chosen. Before a retention choice, partition totals are not computed. Incomplete identities are never collapsed. Source references count the header as record 1; multiline quoted cells do not create new records. Only the first parser error is reported; parser physical lines may differ from logical records. Failed parsing produces no partial review or clean-file claim. No numeric conversion, repair, semantic validation or import certification.
File contents are processed locally without uploads, page-URL input values, analytics, advertising or browser storage. A host serves pages/assets. Downloads save new files; Clear removes displayed input and reports but does not delete saved downloads. Treat exact JSON and literal CSV as sensitive original values.