Remove Duplicate Lines

Paste a list and every repeat is marked in place with a count of how many times it appears, before anything is dropped. Decide for yourself whether case and spacing count as a difference, and keep the first, the last, the unique or only what repeats.

Your list0 chars
Deduplicated
Your result appears herePaste text on the left. It is processed in your browser and never sent anywhere.
01 · The hard part

Deciding what counts as the same line.

Removing duplicates is trivial. Deciding what a duplicate is, is not, and it is the entire job. Two lines are identical or they are not, but the useful question is whether two lines that differ slightly are the same entry — and that depends on what the list is for.

An email list wants case ignored, because a mailbox does not care about capitalisation and sending the same person two copies is a real cost. A log file wants case respected, because case is data. A pasted list wants trailing spaces ignored, because they are invisible and arrived by accident. A file of case-sensitive identifiers wants them respected, because merging two distinct records is worse than leaving a duplicate in.

So the comparison is yours to set, and every line that will be dropped is marked in your own text first, with a count of how many times it occurs. A deduplicator that quietly removes the wrong rows is worse than not running one, because you will not find out until the totals are wrong.

02 · By job

What to set, depending on what the list is.

  • Email and contact listsDuplicates here cost money and reputation: sending the same message twice raises complaint rates and can trip spam filtering. Addresses are case-insensitive by convention, so Alice@example.com and alice@example.com are the same mailbox and should collapse into one. Turn ignore case on, which is the default.
  • Keyword and tag listsExports from keyword tools overlap heavily, and the same term arrives from three sources with different capitalisation and stray spacing. Deduplicating with case and spacing ignored gives a true count of distinct terms, which is the number you actually wanted.
  • Log filesThe opposite settings. Case is meaningful in a log, and so is whitespace, so both should be respected rather than ignored. Switch the mode to show only the repeated lines and a hundred thousand lines of log collapse into the handful of distinct errors that are actually recurring.
  • CSV and data preparationDuplicate rows inflate every aggregate downstream, and they are hard to see in a spreadsheet of any size. Deduplicating before import is far easier than reconciling totals afterwards. Watch case here: identifiers are often case-sensitive, so ignore case may merge rows that are genuinely distinct.
  • URL and link listsCrawls and exports produce the same URL repeatedly. Note that a trailing slash makes two URLs different strings even though they resolve to the same page, so deduplication catches exact repeats and leaves near-duplicates for you to spot.
03 · The modes

Four questions you can ask of a list.

Keep the first of each

The default. Every distinct line survives once, at the position where it first appeared, so the original order is preserved. This is what you want for almost every list.

Keep the last of each

Same result, different position: each distinct line survives at its final occurrence instead of its first. Useful when later entries are more recent and you want the most current version of each record kept.

Keep only lines that appear once

Drops every line that was repeated at all, including its first occurrence. This is not deduplication so much as isolation: it answers the question of which entries are unique to this list, which is how you find the rows one source has and another does not.

Show only the repeated lines

The inverse, and the one people forget exists. It reports exactly what was duplicated and how many times, which is the diagnostic view. Use it before you deduplicate anything important, so you can see what is about to be dropped.

Ignore case and spacing

These decide what counts as the same line. Ignoring case merges Alice and alice; ignoring spacing merges a line with a trailing space and one without. Both are right for email and keyword lists and wrong for logs and case-sensitive identifiers, which is why they are separate switches rather than one setting.

04 · Before you deduplicate

Clean the lines first, or the comparison lies.

Deduplication compares strings, and a string carrying an invisible difference is a different string. This is the most common reason a list still has duplicates in it after being deduplicated: the two entries were never equal to begin with.

Trailing spaces are the usual culprit, which is why ignore spacing exists here. But a non-breaking space, a zero-width character or a byte-order mark will also make two identical-looking lines compare as different, and no amount of trimming will find those. Run the zero-width space remover over a list first if entries that look the same refuse to collapse.

Lists pasted from documents or AI output often need the AI space remover too, and a list whose records span multiple lines needs the line breaks resolved before any line-based comparison will mean anything.

Lists generated by an assistant are the worst case, because they arrive with hidden characters, irregular spacing and inconsistent capitalisation together. The ChatGPT text cleaner resolves the lot in one pass, and the ChatGPT watermark remover tells you which invisible characters were responsible.

05 · FAQ

Duplicate line questions.

How do I remove duplicate lines from a list?

Paste it above. Every duplicate is marked in your own text with a count of how many times it appears, and the first occurrence of each line is kept so the original order survives. Nothing is dropped before you have seen what is going.

Does it keep the original order?

Yes, by default. The first occurrence of each distinct line stays exactly where it was, so a list you had already ordered stays ordered. Sorting alphabetically is available as a separate option if you want it.

Is it case sensitive?

It is your choice, and the right answer depends on the job. Ignore case is on by default because the most common use is email and keyword lists, where Alice and alice are the same thing. Turn it off for logs, identifiers and anything where capitalisation carries meaning.

What does ignore spacing do?

It trims each line before comparing, so a line with a trailing space is treated as identical to the same line without one. Trailing spaces are invisible and extremely common in pasted lists, and without this option they make two identical-looking lines count as different.

How do I find duplicates without removing them?

Set keep to show only the repeated lines. You get a list of exactly what was duplicated, which is the diagnostic view rather than the cleanup one. It is worth running before you deduplicate anything you cannot easily reproduce.

How do I find the lines that appear only once?

Set keep to only lines that appear once. That drops every repeated line entirely, including its first occurrence, leaving the entries that are unique to the list. This is how you compare two exports and find what one has that the other does not.

How do I remove duplicates in Excel?

The Data tab has a Remove Duplicates button that works on selected columns. It is genuinely good, with one caveat worth knowing: it compares raw cell values, so a trailing space makes two entries distinct and they will both survive. Running TRIM over the column first avoids that.

How do I remove duplicates in Google Sheets?

Either Data then Data cleanup then Remove duplicates, or the UNIQUE function in a formula, which has the advantage of leaving the original data intact. Both are exact comparisons, so trailing spaces and capitalisation differences will keep entries that look identical.

How do I remove duplicate lines in Notepad++?

There is a built-in command under the TextFX menu in older builds, and a Line Operations submenu in current ones with Remove Duplicate Lines. It compares exactly and preserves order. It does not offer a case-insensitive mode, which is the usual reason people end up somewhere else.

Does it remove blank lines?

By default yes, since a list of a thousand entries with scattered empty lines almost never wants them treated as duplicates of each other. Turn drop blank lines off if the empty lines are meaningful separators you want kept.

Can I sort the result alphabetically?

Yes, with the sort option. It is off by default because deduplication and sorting are separate decisions, and a list that arrived in a deliberate order should not be reordered as a side effect of removing duplicates.

Is there a limit on list size?

No. Nothing is uploaded, so there is no request size to exceed. Very large lists may take a moment to render the marked view, since every duplicate is highlighted individually.

Will it work on a CSV?

It deduplicates whole lines, so it works on a CSV where each row is one line. Be careful with files where a field contains a line break, since those records span multiple lines and will not compare as you expect. Those need the line breaks resolved first.

Is my data uploaded anywhere?

No. Everything runs in your browser. Nothing is transmitted, nothing is stored, and closing the tab discards it, which matters when the list is customer addresses.