← Back to Dev Blog

Building an Automated, Searchable Email Backup System for Mozilla Thunderbird

Email is the backbone of our digital identity. Invoices, contracts, travel tickets, and decades of personal correspondence all sit inside our inboxes. Yet, despite its critical importance, most people treat email backups as an afterthought—often relying on proprietary exports or opaque, monolithic archive files.

In this post, we will explore why file-level .eml backups are superior for long-term data preservation, and how to build a fully automated, incremental daily backup pipeline using Python, a robust Windows Batch wrapper with Conda support, and the Windows Task Scheduler.


1. Why Searchable, File-Based Email Backups Matter

The Problem with Monolithic Archives (.pst, .mbox, .zip)

By default, clients like Mozilla Thunderbird store thousands of emails inside large, monolithic text files known as MBOX format (or .pst in Outlook). While efficient for the mail client, these massive database files have severe drawbacks for backups:

  1. Corruption Risk: If a single byte inside a 15 GB MBOX file gets corrupted during a disk sync, you risk losing thousands of messages.
  2. Terrible Backup Efficiency: Even if you only received one single newsletter today, your backup tool (Veeam, Borg, Restic, OneDrive) has to re-read and re-deduplicate a massive multi-gigabyte file.
  3. No Direct Searchability: You cannot inspect a message without loading the entire archive back into an email client.
  4. Vendor Lock-in: Opening an old archive five or ten years from now might require legacy software that is no longer maintained.

The Superiority of .eml (RFC 822)

An .eml file is an open, plaintext representation of a single email message including headers, body, and attachments.

  • Universal Compatibility: An .eml file can be opened by Windows Mail, Apple Mail, Outlook, Thunderbird, or even Notepad.
  • Instant OS Search: Windows Search, macOS Spotlight, or command-line tools like grep and ripgrep can index and search through your messages in milliseconds.
  • True Incremental Backups: When each email is its own file, backing up 50 new emails today means writing exactly 50 small files. Your snapshot backups take seconds rather than hours.

2. The Extraction Engine: Python

Thunderbird’s profile folder contains account configurations, indices (.msf), and actual mail data inside Mail/ (POP3 & Local Folders) and ImapMail/ (IMAP accounts). Subfolders are organized using directories ending in .sbd.

To extract these messages reliably, we wrote a standalone Python script that:

  1. Intelligently locates the active profile: Modern Thunderbird installations use a dedicated profile defined in [Install<HASH>] inside profiles.ini (e.g., *.default-release), while older profiles (*.default) might just sit empty. The script evaluates profile disk usage to guarantee it targets the profile that actually holds your data.
  2. Recreates the folder hierarchy: It strips Thunderbird's internal .sbd directory naming to produce a clean, human-readable directory tree.
  3. Guarantees deduplication: Every message is saved with a deterministic filename based on its Message-ID hash. If a message has already been backed up, it is skipped instantly.
src/thunderbird_backup.py
import os
import re
import sys
import hashlib
import mailbox
import configparser
from pathlib import Path
from email.header import decode_header, make_header

# ========================================================
# CONFIGURATION
# Set your preferred local or external backup directory
# ========================================================
BACKUP_DESTINATION = Path(r"C:\Backups\Thunderbird_EML")
# ========================================================

def ensure_destination_dir(dest: Path):
    """Ensure the destination directory exists prior to processing."""
    try:
        dest.mkdir(parents=True, exist_ok=True)
        print(f"[OK] Destination ready: {dest}")
    except Exception as e:
        print(f"[ERROR] Could not create destination {dest}: {e}")
        sys.exit(1)

def get_profile_mail_size(profile_path: Path) -> int:
    """Calculate the total size of mail data to evaluate profile activity."""
    total_size = 0
    for folder_name in ("Mail", "ImapMail"):
        mail_dir = profile_path / folder_name
        if mail_dir.exists():
            for root, _, files in os.walk(mail_dir):
                for f in files:
                    if not f.endswith(".msf"):
                        try:
                            total_size += (Path(root) / f).stat().st_size
                        except OSError:
                            pass
    return total_size

def find_best_thunderbird_profile() -> Path:
    """Locate the active profile containing actual mail data."""
    appdata = os.environ.get("APPDATA")
    tb_base = Path(appdata) / "Thunderbird"
    profiles_ini = tb_base / "profiles.ini"

    if not profiles_ini.exists():
        raise FileNotFoundError(f"profiles.ini not found at {tb_base}")

    config = configparser.ConfigParser()
    config.read(profiles_ini, encoding="utf-8")

    candidate_paths = []

    # 1. Check modern install-based defaults ([Install...])
    for section in config.sections():
        if section.startswith("Install") and config.has_option(section, "Default"):
            raw_path = config.get(section, "Default")
            p = tb_base / raw_path
            if p.exists():
                candidate_paths.append(p)

    # 2. Check traditional profile sections ([Profile...])
    for section in config.sections():
        if section.startswith("Profile") and config.has_option(section, "Path"):
            raw_path = config.get(section, "Path")
            is_rel = config.get(section, "IsRelative", fallback="1") == "1"
            p = (tb_base / raw_path) if is_rel else Path(raw_path)
            if p.exists() and p not in candidate_paths:
                candidate_paths.append(p)

    # Fallback to direct directory scan if profiles.ini is unusual
    profiles_dir = tb_base / "Profiles"
    if profiles_dir.exists():
        for p in profiles_dir.iterdir():
            if p.is_dir() and p not in candidate_paths:
                candidate_paths.append(p)

    print(f"[INFO] Discovered profile candidates ({len(candidate_paths)}):")
    best_profile = None
    max_size = -1

    for p in candidate_paths:
        size = get_profile_mail_size(p)
        mb = size / (1024 * 1024)
        print(f"       - {p.name} ({mb:.1f} MB mail data)")
        if size > max_size:
            max_size = size
            best_profile = p

    return best_profile

def sanitize_filename(name: str, max_len: int = 60) -> str:
    name = re.sub(r'[\\/*?:"<>|]', '_', name)
    name = re.sub(r'\s+', ' ', name).strip()
    return name[:max_len] if name else "No_Subject"

def parse_header_text(header_val: str) -> str:
    if not header_val:
        return ""
    try:
        return str(make_header(decode_header(header_val)))
    except Exception:
        return str(header_val)

def process_mbox(mbox_path: Path, dest_dir: Path):
    dest_dir.mkdir(parents=True, exist_ok=True)
    try:
        box = mailbox.mbox(mbox_path)
    except Exception as e:
        print(f"[WARN] Failed to open MBOX {mbox_path.name}: {e}")
        return 0, 0

    new_count = 0
    skip_count = 0

    for msg in box:
        msg_id = msg.get("Message-ID", "").strip()
        if not msg_id:
            msg_id = hashlib.sha256(msg.as_bytes()[:1024]).hexdigest()

        id_hash = hashlib.md5(msg_id.encode('utf-8', errors='ignore')).hexdigest()[:10]
        subject = parse_header_text(msg.get("Subject", "No_Subject"))
        filename = f"{id_hash} - {sanitize_filename(subject)}.eml"
        file_path = dest_dir / filename

        # Incremental skip: already backed up
        if file_path.exists():
            skip_count += 1
            continue

        try:
            with open(file_path, "wb") as f:
                f.write(msg.as_bytes())
            new_count += 1
        except Exception as e:
            print(f"[ERROR] Could not write {filename}: {e}")

    return new_count, skip_count

def run_backup():
    ensure_destination_dir(BACKUP_DESTINATION)

    profile_dir = find_best_thunderbird_profile()
    print(f"\n[INFO] Selected Active Profile: {profile_dir}")
    print(f"[INFO] Backup Destination: {BACKUP_DESTINATION}\n")

    storage_roots = [profile_dir / "Mail", profile_dir / "ImapMail"]
    total_new = 0
    total_skipped = 0
    folders_processed = 0

    ignore_extensions = {".msf", ".dat", ".json", ".sqlite", ".ini", ".txt", ".bak"}

    for root in storage_roots:
        if not root.exists():
            continue

        for dirpath, _, filenames in os.walk(root):
            dirpath = Path(dirpath)
            rel_path = dirpath.relative_to(root)
            clean_parts = [p.removesuffix(".sbd") for p in rel_path.parts]
            folder_target_dir = BACKUP_DESTINATION / root.name / Path(*clean_parts)

            for file in filenames:
                file_path = dirpath / file
                if file_path.suffix.lower() in ignore_extensions:
                    continue
                if file_path.stat().st_size == 0:
                    continue

                new_cnt, skip_cnt = process_mbox(file_path, folder_target_dir / file)
                if (new_cnt + skip_cnt) > 0:
                    folders_processed += 1
                    print(f"  -> Folder '{file}': +{new_cnt} new, {skip_cnt} existing.")
                    total_new += new_cnt
                    total_skipped += skip_cnt

    print("\n----------------- SUMMARY -----------------")
    print(f"Mail Folders Processed : {folders_processed}")
    print(f"New Emails Exported    : {total_new}")
    print(f"Existing Emails Skipped: {total_skipped}")
    print("-------------------------------------------")

    if (total_new + total_skipped) == 0:
        print("[ERROR] 0 emails found! Check if 'Offline Synchronization' is enabled for IMAP.")
        sys.exit(2)

if __name__ == "__main__":
    run_backup()

3. The Windows Conda Wrapper: run_backup.cmd

Many developers run Python inside Anaconda or Miniconda environments. However, Conda is typically not added to the Windows system PATH by default to prevent DLL collisions with system tools.

Furthermore, executing batch files in Windows comes with subtle traps:

  • Wrapping a command in double quotes (e.g., call "conda") causes cmd.exe to stop searching the system PATH, triggering the error: The system cannot find the path specified.
  • Hardcoding paths makes the script unportable.

To solve this, our batch script dynamically identifies the absolute location of conda.bat via where.exe, uses %~dp0 to operate relative to its own folder, writes to a local backup.log, and executes cleanly via conda run:

scripts/run_backup.cmd
@echo off
setlocal EnableExtensions EnableDelayedExpansion

:: 1. Derive script directory dynamically
set "BASE_DIR=%~dp0"
set "SCRIPT_PATH=%BASE_DIR%thunderbird_backup.py"
set "LOG_PATH=%BASE_DIR%backup.log"
set "CONDA_ENV_NAME=base"

echo ======================================================== >> "%LOG_PATH%"
echo [START] Backup started on %date% at %time% >> "%LOG_PATH%"

:: 2. Resolve absolute path to Conda
set "CONDA_CMD="
for /f "delims=" %%I in ('where.exe conda.bat 2^>nul') do (
    set "CONDA_CMD=%%~fI"
    goto :FOUND_CONDA
)
for /f "delims=" %%I in ('where.exe conda.exe 2^>nul') do (
    set "CONDA_CMD=%%~fI"
    goto :FOUND_CONDA
)

:: Fallback: search common Conda install locations
for %%P in (
    "%USERPROFILE%\miniconda3\condabin\conda.bat"
    "%USERPROFILE%\anaconda3\condabin\conda.bat"
    "%LOCALAPPDATA%\miniconda3\condabin\conda.bat"
    "%LOCALAPPDATA%\anaconda3\condabin\conda.bat"
    "C:\ProgramData\miniconda3\condabin\conda.bat"
    "C:\ProgramData\anaconda3\condabin\conda.bat"
) do (
    if exist %%P (
        set "CONDA_CMD=%%~fP"
        goto :FOUND_CONDA
    )
)

echo [ERROR] Conda executable could not be located! >> "%LOG_PATH%"
exit /b 1

:FOUND_CONDA
echo [INFO] Conda located at: %CONDA_CMD% >> "%LOG_PATH%"

:: 3. Validate Python script presence
if not exist "%SCRIPT_PATH%" (
    echo [ERROR] Python script not found at: %SCRIPT_PATH% >> "%LOG_PATH%"
    exit /b 1
)

:: 4. Execute inside Conda Environment
echo [INFO] Running Python inside Conda environment '%CONDA_ENV_NAME%'... >> "%LOG_PATH%"
call "%CONDA_CMD%" run -n "%CONDA_ENV_NAME%" python "%SCRIPT_PATH%" >> "%LOG_PATH%" 2>&1

if %errorlevel% neq 0 (
    echo [ERROR] Backup execution failed with error code %errorlevel%. >> "%LOG_PATH%"
    exit /b %errorlevel%
)

echo [SUCCESS] Backup completed successfully at %time%. >> "%LOG_PATH%"
exit /b 0

4. Automation: Set-and-Forget Daily Runs

With the Python script and the CMD wrapper living in the same directory (e.g., C:\Scripts\ThunderbirdBackup), you can automate the process using Windows Task Scheduler.

Run the following command in an administrative Command Prompt:

cmd
schtasks /create /tn "DailyThunderbirdBackup" /tr "C:\Scripts\ThunderbirdBackup\run_backup.cmd" /sc daily /st 21:00

Pro Tip: Preventing Command Window Popups

To prevent a black terminal window from flashing onto your screen every evening, create a one-line VBScript file named silent_run.vbs:

vbscript
Set WshShell = CreateObject("WScript.Shell")
WshShell.Run "C:\Scripts\ThunderbirdBackup\run_backup.cmd", 0, False

Then, point Task Scheduler to execute wscript.exe "C:\Scripts\ThunderbirdBackup\silent_run.vbs". It will run completely invisible in the background.


5. Verification & Log Output

When the script runs, check backup.log to see a detailed summary:

text
======================================================== 
[START] Backup started on 2026-03-15 at 21:00:01,12 
[INFO] Conda located at: C:\Users\<Username>\Miniconda3\Library\bin\conda.bat 
[INFO] Running Python inside Conda environment 'base'... 
[OK] Destination ready: C:\Backups\Thunderbird_EML
[INFO] Discovered profile candidates (2):
       - <legacy_hash>.default (0.0 MB mail data)
       - <active_hash>.default-release (4120.5 MB mail data)

[INFO] Selected Active Profile: C:\Users\<Username>\AppData\Roaming\Thunderbird\Profiles\<active_hash>.default-release
[INFO] Backup Destination: C:\Backups\Thunderbird_EML

  -> Folder 'INBOX': +4 new, 18420 existing.
  -> Folder 'Sent': +1 new, 4310 existing.

----------------- SUMMARY -----------------
Mail Folders Processed : 12
New Emails Exported    : 5
Existing Emails Skipped: 22730
-------------------------------------------
[SUCCESS] Backup completed successfully at 21:00:04,45.

Notice that inspecting over 22,000 messages and exporting the 5 newest ones took under 4 seconds.


Summary and Key Takeaways

  1. Granular .eml Preservation: By breaking away from monolithic mail databases and treating emails as individual .eml files, backups become resilient against single-file corruption.
  2. Instant Searchability: Standard OS indexing tools (Windows Search, Spotlight, ripgrep) can query and locate specific correspondence in milliseconds.
  3. High-Speed Incremental Sync: Daily differential backups take seconds rather than minutes or hours, transferring only new messages.
  4. Resilient Conda Automation: Dynamic executable resolution in batch wrappers guarantees robust unattended Task Scheduler runs.