Spreadsheets

Find duplicate records without deleting good data

Matching names do not necessarily identify the same record.

OfficeCubs · · 2 min read

Find duplicate records without deleting good data

Matching names do not necessarily identify the same record. Produce candidates with reasons before removing anything from a customer export.

Prepare the workbook

Keep an unchanged source copy and explain column meanings, units and date boundaries. Use a spreadsheet or data role with compatible local tools, or the prepared sandbox. Ask for a rejected-row report instead of silently deleting inconvenient records. OfficeCubs text previews do not recalculate formulas; open the final workbook in a spreadsheet application and check representative formulas and totals.

Inputs for this workflow: customers.csv with stable customer IDs and optional email addresses.

Work through the task

  1. Define exact matches separately from possible matches.
  2. Group candidates and retain every source row identifier.

A brief you can adapt

Write duplicate-candidates.csv with match_rule, source_ids and suggested_action. Compare normalized email only when present. Do not merge records or overwrite customers.csv.

The filenames above are examples. Replace them with your actual inputs and destination before submitting the task.

Review the deliverable

Inspect a repeated name with different emails and a shared company inbox used by different people.

Where this approach can fail

A fuzzy similarity score is a review aid, not permission to merge identities.

For the related desktop setup, see missing document tools.

Keep exploring

Setup and troubleshooting in Help