Data · THE NO-PANIC PLAN
Clean a CSV file for a reliable data import
Prepare a CSV import without losing records or changing values silently. Preserve the source, document each transformation, and check the result against the receiving system?s field contract before loading it.
MISSION clean csv data
Start this workflowTHE REAL-WORLD BIT
What happens outside this browser tab?
Freeze a source copy, profile the table and destination contract, reconcile headers and candidate duplicates, apply logged reversible changes, then validate the saved output and reconcile its records.
YOUR CHECKLIST, WITH FEWER DRAMATIC SIGHES
One step at a time.
Follow the order below. If a step names a Nirmion tool, its link is right there with it.
- 01
Freeze the source and define the import contract
Keep an untouched copy with source, owner, export time, row count, encoding and intended destination recorded. Obtain the destination?s required headers, field types, key rules, null behavior, date/number locale, delimiter, line ending and error format. Confirm you are authorized to handle the data; work on a representative masked sample if it contains personal, customer, health or financial information. Do not treat a successful CSV parse as proof that the destination accepts its values.
- 02
Profile the CSV structure before editing
Open the file with a CSV-aware parser, not line splitting. Confirm whether a header is present; record delimiter, encoding and line-ending behavior; inspect column names, field counts, row count, empty cells, candidate keys, and representative quoted fields. RFC 4180 describes common conventions: fields with commas, quotes or line breaks need quoting, embedded quotes are doubled, field counts should be consistent, and spaces are part of a field. It is informational and implementations vary, so use the receiving parser?s documented dialect when it differs.
- 03
Resolve headers and review duplicates without deleting blindly
Map each source header to the exact destination field and preserve the mapping log. Nirmion CSV Column Mapper can rename selected columns using an explicit source-to-target mapping. Define record identity with the data owner before checking duplicates; use CSV Duplicate Finder on a safe sample to find rows matching the selected fields and retain source row numbers. Review each candidate against the identity rule before any merge or deletion: equal names or addresses alone may not mean the same entity.
- 04
Apply approved, traceable transformations
Make changes in a separate working copy using a CSV-aware editor or transformation pipeline. Document the rule, affected column, input/output counts and exceptions for every trim, normalization, type conversion, null substitution, split, merge or correction. Preserve leading zeroes in identifiers and source precision; do not guess ambiguous dates, currencies, missing values or entity matches. Keep rejected and unresolved rows in a separate exception file for owner resolution.
- 05
Validate the export and reconcile it to the source
Save a new output and reopen it with the same parser and declared dialect. Check the required header names/order, consistent field counts, parseable types, required/non-null fields, key uniqueness, exception count and agreed totals. Validate columns against receiver-provided metadata/schema where available; W3C?s tabular model describes using metadata such as datatypes to validate tabular values, while the actual importer contract remains authoritative. Reconcile input rows to accepted plus explicitly rejected rows, test a small approved sample in a non-production destination, and retain the original, mapping, change log and validation result.
THE HELPER CREW
Tools for the fiddly bits.
These are the currently published Nirmion tools matched to this guide. Open a tool page for its accepted inputs and limits.
RECEIPTS, PLEASE
Sources & review notes
Each source is linked to the steps it supports. Open it to check its scope and current guidance.
Source checked 2026-10-04