Nirmion
HelpLog in Find a tool

Data · THE NO-PANIC PLAN

Deduplicate a CSV without losing records

Use this workflow when a CSV export contains repeated records and you need a defensible way to identify and remove them. It distinguishes exact row matches from records matching only on selected columns, explains how to review groups before deletion, and shows where Nirmion's browser-local CSV Duplicate Finder fits. It identifies duplicates; it does not choose the correct record or delete rows for you.

MISSION Find and remove duplicate CSV records using a documented key and survivor rule, while preserving the original file, quoted fields, an audit trail, and a verified output.

Review duplicate groups with CSV Duplicate Finder

THE REAL-WORLD BIT

What happens outside this browser tab?

Protect the original, confirm the CSV structure, define the exact key and which record should survive, preview duplicate groups with source row numbers, remove only approved rows in a working copy, then reopen and reconcile the output against the original and your decision log.

YOUR CHECKLIST, WITH FEWER DRAMATIC SIGHES

One step at a time.

Follow the order below. If a step names a Nirmion tool, its link is right there with it.

  1. 01

    Make a protected working copy

    Save an untouched copy of the export and work on a separately named copy. Record who owns the data, why deduplication is authorized, where the file came from, its export date, expected row count, and where the cleaned file will be stored. Limit access to the people who need it. If the CSV contains personal, financial, health, confidential, or otherwise restricted data, follow your organization’s approved handling rules; do not paste it into an online tool unless that use is expressly approved. Keep the original available for rollback because spreadsheet removal commands can permanently delete rows from the working sheet.

  2. 02

    Confirm the CSV parses as the intended table

    Open the copy with the application or parser that will receive the cleaned result. Confirm the delimiter, character encoding, header row, column names, and expected field count; compare a few records with the source system. Do not split the file on raw newline characters: a quoted CSV field can itself contain commas, quote marks, or line breaks. RFC 4180 documents these common conventions but is informational and notes that CSV implementations differ, so confirm the actual import preview rather than assuming every file follows one universal format. Stop if records shift columns, quotes are malformed, or identifiers such as 00123 are converted to 123.

  3. 03

    Define what counts as the same record

    Write the matching key before removing anything. Use a stable unique identifier when one exists. Otherwise choose the smallest justified set of columns that identifies the same real-world record, such as an order ID plus line number; matching on a name alone can merge different people. Decide which record survives using an auditable rule, such as the authoritative system, latest verified update, or a data owner’s decision. Specify whether case, leading/trailing spaces, blank values, punctuation, and dates are considered different. Do not normalize values until you have checked that the transformation is valid for this dataset.

  4. 04

    Preview duplicate groups and record the decision

    Use Nirmion’s CSV Duplicate Finder with only data you are authorized to process. It runs in the browser, compares exact cell values in the selected column or columns, and returns the duplicate group, original CSV row number, and full row for every member of a repeated group. Select the same columns you documented as the key; selecting none compares the full row. It does not trim spaces, ignore letter case, select a survivor, or delete records. Review every returned group against the source system and your survivor rule, then keep a private decision log of rows approved for removal and any ambiguous groups. In Power Query, select the key columns and choose Home > Remove Rows > Remove Duplicates only after that review; preserve the untouched source because Microsoft warns that removal changes the result.

  5. 05

    Remove approved rows and reconcile the export

    Remove only reviewed duplicate rows from the working copy, following the documented survivor rule; leave unresolved groups untouched and ask the data owner. Save to a new output path. Reopen the exported CSV using the intended receiving application and confirm headers, field counts, encoding, dates, leading-zero identifiers, and quoted commas or line breaks still parse correctly. Re-run the same-key duplicate check and confirm no approved duplicates remain. Reconcile input rows minus approved removals to the output row count, check unique identifiers and key totals, and inspect representative kept and removed groups against the original. Record the output name, method, date, reviewer, row counts, exceptions, and rollback location before marking the task complete.

THE HELPER CREW

Tools for the fiddly bits.

These are the currently published Nirmion tools matched to this guide. Open a tool page for its accepted inputs and limits.

RECEIPTS, PLEASE

Sources & review notes

Each source is linked to the steps it supports. Open it to check its scope and current guidance.

Source checked 2026-10-05