FIRST CH TOOLS / Text / 63 LINE TOOLS

Text Line Tools: Remove Duplicates, Sort & Number Lines

Tidy a list with one item per line: remove duplicate lines, sort, shuffle, number the lines, drop blank lines, trim spaces and add a prefix or suffix to every line. Only the steps you tick are applied, once each and always in the same order, so "trim, then remove duplicates, then sort, then number" is a single pass. Sorting compares numbers by value: item2 comes before item10, and -5 before 1.5.

Ticked steps run from top to bottom

Common combinations
01Spaces
02Blank lines
03Duplicates
04Order
05Prefix & suffix
06Line numbers
Text to process
Result
0
Lines in
0
Lines out
0
Duplicates removed
0
Blank lines removed
Checks

    This tool works line by line. Character-level conversion such as making full-width ABC and half-width ABC match belongs to the Full-width ⇄ Half-width Converter; to see how two lists differ, use the Diff Checker.

    How to Use

    1. Paste the listOne item per line in the left-hand box — a column copied out of a spreadsheet or a list of URLs pulled from a log works as it is.
    2. Tick the stepsChoose any of 01–06. Buttons such as "Dedupe & sort" switch on a common combination in one click. The result on the right updates with every change.
    3. CopyTake the right-hand box back to where the list came from. The lines that were duplicated, with their counts, and any near-duplicates that slipped through are listed under Checks.

    Order of operations and sorting rules

    Steps run 01 → 06, once each

    Whatever order you tick them in, the steps always run spaces → blank lines → duplicates → order → prefix & suffix → line numbers. Trimming comes first, so "apple" and "apple␣" are removed as duplicates; numbering comes last, so the numbers run 1, 2, 3… in the sorted order. With both a prefix and numbers, the number goes outside: 1. - apple.

    Number-only lines sort by value; other lines compare their digit runs as numbers

    With "compare numbers by value" on, lines that are just a number (-10, 1.25, 1,000, full-width 12) are ordered by value and placed before lines containing text. Other lines treat each run of digits as one number, giving item1 → item2 → item10. Switch it off and lines are compared character by character: item1 → item10 → item2.

    Case, kana and kanji

    Sorting uses the browser's Japanese collation (Intl.Collator), which orders Latin letters the way an English dictionary does:abc, Abc and ABC sit next to each other (ties that collation cannot separate are broken by code point, so the same input always gives the same output). Hiragana and katakana are interleaved in gojūon order. Kanji carry no reading, so they are not sorted by pronunciation but in JIS level-1 order — roughly the order of a representative on'yomi reading — for example 札幌 → 大阪 → 東京 → 名古屋. For true reading order, put a kana reading at the start of each line first.

    Three ways to handle duplicates

    Keep the first occurrence is the usual dedupe: one copy stays where it first appeared. Drop every line that repeats keeps only the lines that occur once — stack two lists on top of each other first and this gives you the items that are in only one of them. Keep only the repeated lines, once each is the opposite: on two stacked lists it gives you the items that are in both.

    About This Tool

    Tidying a list is the same few chores over and over: remove duplicate email addresses, re-sort a column of product codes copied out of a spreadsheet, turn a list of URLs into Markdown bullets, shuffle a draw. A spreadsheet or an editor can do each of them, but trimming, deduplicating and sorting are separate jobs every time. Here you tick the steps you need and they run together, in a fixed order.

    Duplicates that "survive" usually differ in ways you cannot see. A trailing space, an ideographic space, Apple versus apple, full-width A versus half-width A — they look the same on screen and are still different lines. Checks reports how many groups of near-duplicates that differ only in surrounding spaces, letter case or width are left. The first two are settings on this page; for width, run the text through the Full-width ⇄ Half-width Converter and come back.

    Shuffling is Fisher–Yates, drawing every step from crypto.getRandomValues with rejection sampling. Unlike stretching a single seed into a pseudo-random stream, no possible order is left out however long the list. Pressing "Shuffle again" draws fresh random numbers. Adding seed= to the URL reproduces the same order from the same input, but that reproducible mode builds the order from a 32-bit seed, so with 13 or more lines it cannot reach every possible order — leave seed= off for a draw. The popular shortcut of calling sort() with Math.random() is known to give a biased order.

    A trailing newline is kept. If the pasted text ends with a newline, so does the result. As for line endings, the browser's text box turns everything into LF the moment you paste, so the result is LF too. When you need CRLF for Windows, run the result through the Encoding Converter.

    Nothing you paste ever leaves the browser — every step runs on your machine, and a list of 100,000 lines is processed in a moment. On this page the site analytics are not given the query part of the URL (anything after ?). Directly callable via URL parameters: /en/lines/?text=b%0Aa%0Ab&dedupe=1&sort=asc / /en/lines/?preset=numbered&sep=paren&pad=1

    Other Tools