Skip to main content

Command Palette

Search for a command to run...

Stop Using Heavy Libraries: How I Cleaned and Deduplicated Text Files with 25 Lines of Pure Python

Updated
•4 min read•View as Markdown
V
Exploring technology, building projects, and writing about it. Aspiring technical writer dedicated to making technical documentation accessible for everyone

In my first year of Computer Science, everyone kept telling me, "If you want to clean data in Python, you have to use Pandas."

But over my summer break, I wanted to challenge myself. What if I had a messy text file full of names and emails, and I wanted to clean it up using nothing but raw, pure Python. No heavy external libraries, no massive installations just standard library logic sounds interesting right?

I sat down and wrote a 25 line script that parses a dirty data file, handles missing commas, standardizes text capitalization, and eliminates duplicates instantly. Well now, you can do that too!

Here is the exact code, how it works under the hood, and the logic jumps I had to figure out to make it bulletproof.

The Problem: The Anatomy of "Dirty Data"

Imagine you have a file called dirty_data.txt. It contains user sign-ups, but it's full of human errors:

  • Some lines are missing commas entirely.

  • People typed their names in all lowercase (john doe) or messy caps (jOhN dOe).

  • The same email address signed up multiple times.

Overcoming Edge Cases:

Honestly, I didn't write my final code just like that magically in the first attempt.

When I was writing the logic, I first created a code block to strip out whitespaces. Then, I wrote a block to split each line at the comma into two parts, "name" and "email".

But then I thought, what if a line doesn't contain a comma at all? Initially, I thought line.split(",", 1) might just handle it smoothly, but I quickly realized a major catch. If a comma is missing, Python can't split the line into two separate pieces. When the script tries to force that single piece of text into two variables (name, email), it throws a crash ValueError: not enough values to unpack.

To build a safety net for this, I wrapped the splitting logic in a try and except ValueError block. Now, instead of crashing the entire script, my code gracefully catches the unpacking error, prints a warning message, logs the malformed line, and jumps straight to the next row using continue.

# maxsplit=1 prevents breaking if a name contains a comma
        try:
            name, email = line.split(",", 1)  
        except ValueError:
            print(f"Skipping malformed data on line {line_num}: '{line}'")
            invalid_lines_count += 1
            continue

I used standard string methods to capitalize and lower the letters in the name wherever needed. I made sure that if the same email appears later in the file, the not in check above will catch it and evaluate to False.

# Deduplicate based on email
        if clean_email not in seen_emails:
            seen_emails.add(clean_email)
            clean_file.write(f"{clean_name},{clean_email}\n")

That's how i printed my clean text along with a count of invalid lines.

The Code: My Pure Python Data Janitor

Here is the exact code I wrote to handle this task:

print("Starting the data cleaning process...")

# Track uniques and bad rows
seen_emails = set()
invalid_lines_count = 0

with (
    open("dirty_data.txt", "r") as dirty_file,
    open("clean_data.txt", "w") as clean_file,
):
    for line_num, line in enumerate(dirty_file, 1):
        line = line.strip()
        if not line:
            continue  

        # maxsplit=1 prevents breaking if a name contains a comma
        try:
            name, email = line.split(",", 1)  
        except ValueError:
            print(f"Skipping malformed data on line {line_num}: '{line}'")
            invalid_lines_count += 1
            continue

        # Standardize formatting
        clean_name = name.strip().title()  
        clean_email = email.strip().lower()

        # Deduplicate based on email
        if clean_email not in seen_emails:
            seen_emails.add(clean_email)
            clean_file.write(f"{clean_name}, {clean_email}\n")

print("\nFinished cleaning.")
print(f"Saved to clean_data.txt ({len(seen_emails)} unique records)")
if invalid_lines_count > 0:
    print(f"Skipped {invalid_lines_count} broken lines.")

Want to try cleaning your own files with this script? Check out the full, ready to run code over on github. If this helped you, feel free to leave a star on the repository!