What the duplicate line remover does
A duplicate line is a line of text that appears more than once in a list. They pile up when you merge two email lists, copy keywords from several reports or paste the same items twice by mistake. This tool reads your text one line at a time and removes the repeats.
There are three ways to use it, chosen with the Keep buttons:
- One copy of each line: removes the repeats and keeps the first copy of every line. This is the usual “remove duplicates” result.
- Lines that appear once: drops every line that has a duplicate, including the first copy. Only the truly unique lines are left.
- Only repeated lines: lists the lines that appeared more than once, each shown one time. Use it to find out what was duplicated.
How to use it
- Paste your list into Your list, one item per line.
- Pick an option under Keep.
- Turn on Ignore case if “Apple” and “apple” should count as the same.
- Leave Ignore spaces at line ends on unless trailing spaces matter to you.
- Copy the result with Copy, or save it with Download.
The message under the tool tells you how many lines were removed and how many are left.
Options in detail
Ignore case compares lines without looking at capital letters. The first version is kept, so if “apple” comes before “Apple”, you get “apple”.
Ignore spaces at line ends trims spaces and tabs at the start and end of each line before comparing. The kept line is not changed. Turn it off if leading spaces are meaningful, as in indented code.
Remove blank lines appears when you keep one copy of each line. It is on by default. Turn it off to keep empty lines where they were, which helps when blank lines separate groups in your list. Blank lines are never treated as duplicates of each other.
Show how many times appears with Only repeated lines. It adds the number of copies in brackets after each line.
Examples
All of these use the sample list on the page: apple, banana, Apple, cherry, banana, a blank line, date, cherry, banana.
| Settings | Result |
|---|---|
| One copy of each line | apple, banana, Apple, cherry, date |
| One copy of each line, Ignore case | apple, banana, cherry, date |
| Lines that appear once | apple, Apple, date |
| Only repeated lines, Show how many times | banana (3), cherry (2) |
| Only repeated lines, Show how many times, Ignore case | apple (2), banana (3), cherry (2) |
The results are shown here separated by commas to save space. In the tool, each item stays on its own line.
With the default settings, the tool reports: “Removed 3 duplicate lines. 5 lines left.”
How it works
The tool splits your text at each line break. It then builds a comparison key for every line: the line itself, trimmed if Ignore spaces at line ends is on and lowercased if Ignore case is on. It counts how often each key appears, then walks through the lines in order and decides which to keep. Because the original lines are kept rather than the keys, your spacing and capitals are not altered.
Tips
- Sort after, not before. Removing duplicates keeps your order. If you want the list sorted as well, paste the result into the alphabetical sorter.
- Clean messy lists first. Lines with double spaces in the middle, such as “New York” and “New York”, do not match. The whitespace remover fixes those.
- One item per line. If your items are separated by commas, put each on its own line first, or use the Commas option in the sorter, which can also remove duplicates.
- Text broken across lines from a PDF should be joined first with the line break remover.
Other ways to remove duplicate lines
Excel: select the column, then choose Data, then Remove Duplicates. In Excel 365 you can also use the formula =UNIQUE(A1:A100).
Google Sheets: use Data, then Data cleanup, then Remove duplicates, or the formula =UNIQUE(A1:A100).
Notepad++: Edit, then Line Operations, then Remove Duplicate Lines.
Command line: sort file.txt | uniq removes duplicates but also sorts the lines. awk '!seen[$0]++' file.txt keeps the original order.
Python: list(dict.fromkeys(lines)) keeps the first copy of each line in order.
For a quick list from an email or a web page, pasting it here is usually faster than opening a spreadsheet.