Automation & Scripting Competent¶
When you'd use this
File system automation, task scheduling, batch processing and OS interaction with Python.
Automate repetitive chores — renaming files, moving data, calling tools — with small scripts that save hours of manual work.
What you'll learn¶
- Manipulate files and folders with
pathlib - Copy, move, archive and clean up with
shutil - Batch-process CSV and JSON data
- Run external programs with
subprocess - Read configuration from environment variables
- Schedule scripts to run automatically
- Turn a script into a reusable CLI tool
- Log what your automation does
Automation is where Python earns its "batteries included" reputation. Most of what follows uses only the standard library — no installs required.
1. File System Automation¶
Find, filter, and operate on files and directories programmatically.
Paths with pathlib¶
pathlib is the modern, object-oriented way to work with paths. Prefer it over the older string-based os.path.
from pathlib import Path
p = Path("reports") / "2026" / "summary.txt" # join with /
print(p.name) # 'summary.txt' — file name
print(p.stem) # 'summary' — name without suffix
print(p.suffix) # '.txt' — extension
print(p.parent) # 'reports/2026' — containing folder
print(p.parts) # ('reports', '2026', 'summary.txt')
Why pathlib over os.path: paths become real objects with methods, the / operator reads naturally, and the same code works on Windows and Linux without worrying about \ vs /.
Reading & writing¶
Read or write a whole file in a single call with read_text/write_text — no open()/close() boilerplate — ideal for small config files, notes, and quick scripts.
from pathlib import Path
f = Path("note.txt")
f.write_text("hello", encoding="utf-8") # write (creates/overwrites)
content = f.read_text(encoding="utf-8") # read whole file
# content -> 'hello'
print(f.exists()) # True
print(f.is_file()) # True
print(f.stat().st_size) # size in bytes
Always pass encoding='utf-8'
Without it, Python uses the platform default, which differs between Windows and Linux and is a common source of "works on my machine" bugs.
Finding files with glob¶
Match files by wildcard pattern instead of listing and filtering by hand — glob for one level, rglob to recurse — the fast way to collect "every .py under here".
from pathlib import Path
base = Path(".")
# All .txt files in this folder
for f in base.glob("*.txt"):
print(f)
# All .py files anywhere below this folder (recursive)
for f in base.rglob("*.py"):
print(f)
glob matches one level; rglob recurses into subfolders. Both return a generator of Path objects.
Real example: organize files by extension¶
A classic automation task — sort a messy download folder into subfolders by file type.
from pathlib import Path
def organize(folder: str) -> dict[str, int]:
"""Move each file into a subfolder named after its extension.
Returns a count of files moved per extension.
"""
base = Path(folder)
moved: dict[str, int] = {}
for item in base.iterdir():
if item.is_file():
ext = item.suffix.lstrip(".") or "no_extension"
target_dir = base / ext
target_dir.mkdir(exist_ok=True) # no error if it exists
item.rename(target_dir / item.name)
moved[ext] = moved.get(ext, 0) + 1
return moved
# result = organize("Downloads")
# -> {'txt': 5, 'pdf': 2, 'jpg': 8}
What's happening: iterdir() lists the folder, suffix.lstrip(".") turns .txt into txt, mkdir(exist_ok=True) creates the target folder safely, and rename() moves the file. The function returns a summary dict so the caller knows what it did.
2. Copy, Move, Archive with shutil¶
High-level file operations and zip/tar archiving in a few calls.
pathlib handles single files well; shutil handles whole trees and archives.
import shutil
# Copy a single file (preserves metadata)
shutil.copy2("report.txt", "backup/report.txt")
# Copy an entire directory tree
shutil.copytree("src", "backup/src")
# Move a file or folder
shutil.move("old/data.csv", "archive/data.csv")
# Delete a whole tree (careful — no undo)
shutil.rmtree("temp_folder")
# Create a .zip archive of a folder → returns the archive path
archive_path = shutil.make_archive("backup-2026", "zip", "src")
# archive_path -> 'backup-2026.zip'
shutil.rmtree is irreversible
It permanently deletes the folder and everything in it — there's no recycle bin. Double-check the path, and consider printing what will be deleted before doing it.
Real example: rotating backups¶
Combine shutil and datetime to make a timestamped zip archive — the core of any "nightly backup" job, where each run produces a uniquely named file.
import shutil
from datetime import datetime
from pathlib import Path
def backup(source: str, backup_root: str) -> str:
"""Create a timestamped zip backup of `source`. Returns the archive path."""
stamp = datetime.now().strftime("%Y%m%d-%H%M%S")
dest = Path(backup_root) / f"backup-{stamp}"
archive = shutil.make_archive(str(dest), "zip", source)
return archive
# backup("project", "backups")
# -> 'backups/backup-20260830-154210.zip'
3. Batch Processing Data¶
Loop over many files/records applying the same transformation.
CSV files¶
csv.DictReader treats each row as a dict keyed by the header — much easier than counting columns.
import csv
from pathlib import Path
def total_scores(path: str) -> int:
"""Sum the 'score' column across every row."""
with Path(path).open(encoding="utf-8") as fh:
reader = csv.DictReader(fh)
return sum(int(row["score"]) for row in reader)
# For a file with rows Alice,90 and Bob,75:
# total_scores("data.csv") -> 165
Writing CSV:
import csv
rows = [
{"name": "Alice", "score": 90},
{"name": "Bob", "score": 75},
]
with open("out.csv", "w", newline="", encoding="utf-8") as fh:
writer = csv.DictWriter(fh, fieldnames=["name", "score"])
writer.writeheader()
writer.writerows(rows)
newline='' on Windows
Always open CSV files with newline="". Without it, Windows inserts an extra blank line between every row.
JSON files¶
Read a JSON file into a Python dict, edit it, and write it back — the usual pattern for updating config or state files that both humans and programs touch.
import json
from pathlib import Path
# Read
data = json.loads(Path("config.json").read_text(encoding="utf-8"))
# Modify
data["version"] = 2
# Write back (indent=2 keeps it human-readable)
Path("config.json").write_text(
json.dumps(data, indent=2), encoding="utf-8"
)
Real example: batch-rename with a counter¶
Rename a whole folder of files to a clean, zero-padded sequence (vacation_001.jpg, …) — a common chore for photos, scans, or exported data that arrives with messy names.
from pathlib import Path
def batch_rename(folder: str, prefix: str) -> list[str]:
"""Rename every file to prefix_001.ext, prefix_002.ext, ...
Returns the list of new names.
"""
base = Path(folder)
files = sorted(f for f in base.iterdir() if f.is_file())
new_names = []
for i, f in enumerate(files, start=1):
new_name = f"{prefix}_{i:03d}{f.suffix}" # 001, 002, ...
f.rename(base / new_name)
new_names.append(new_name)
return new_names
# batch_rename("photos", "vacation")
# -> ['vacation_001.jpg', 'vacation_002.jpg', ...]
The {i:03d} format pads the number to 3 digits with leading zeros, so files sort correctly.
4. Running External Programs with subprocess¶
Call other tools and capture their output from your script.
subprocess.run executes another program and waits for it to finish.
import subprocess
# Run a command, capture its output
result = subprocess.run(
["git", "status", "--short"],
capture_output=True, # capture stdout/stderr
text=True, # decode bytes → str
check=True, # raise if exit code != 0
)
print(result.stdout) # the command's output
print(result.returncode) # 0 on success
Key arguments:
| Argument | What it does |
|---|---|
capture_output=True | Collect stdout and stderr instead of printing them |
text=True | Return strings instead of raw bytes |
check=True | Raise CalledProcessError if the command fails |
Pass a list, not a string — and avoid shell=True
Always pass the command as a list of arguments (["ls", "-la"]). Using shell=True with untrusted input opens you to shell-injection attacks. If a value comes from a user, keep it as a separate list item so it's treated as data, never as a command.
Real example: run a command with a timeout¶
Wrap subprocess.run so a hung or failing command returns a clean error string instead of crashing your script — essential when automating tools that might stall.
import subprocess
def run_safely(cmd: list[str], timeout: int = 30) -> str:
"""Run a command, return its stdout, or an error message."""
try:
result = subprocess.run(
cmd, capture_output=True, text=True,
check=True, timeout=timeout,
)
return result.stdout.strip()
except subprocess.TimeoutExpired:
return f"Timed out after {timeout}s"
except subprocess.CalledProcessError as e:
return f"Failed (exit {e.returncode}): {e.stderr.strip()}"
# run_safely([sys.executable, "-c", "print('hi')"]) -> 'hi'
5. Environment Variables & Configuration¶
Read config from the environment so scripts adapt without code changes.
Never hard-code secrets or environment-specific values. Read them from the environment.
import os
# Read with a fallback default
db_host = os.environ.get("DB_HOST", "localhost")
debug = os.environ.get("DEBUG", "false").lower() == "true"
# Read a required value (raises KeyError if missing)
api_key = os.environ["API_KEY"]
Set them before running (PowerShell):
Keep secrets out of your code
API keys and passwords belong in environment variables or a secrets manager — never in the source file, and never committed to git. See the Secrets Management topic for the full picture.
6. Task Scheduling¶
Run scripts on a schedule with cron/Task Scheduler or Python schedulers.
Two approaches: schedule from inside a long-running Python process, or let the operating system run your script on a timer.
In-process scheduling with sched¶
Schedule callbacks to fire after a delay from within a running Python program using the standard-library sched module — use when the timing is part of a larger app rather than a system-level cron job.
import sched, time
scheduler = sched.scheduler(time.monotonic, time.sleep)
def job(label: str) -> None:
print(f"Running {label} at {time.strftime('%H:%M:%S')}")
# Run `job` 2 seconds and 4 seconds from now
scheduler.enter(2, priority=1, action=job, argument=("first",))
scheduler.enter(4, priority=1, action=job, argument=("second",))
scheduler.run() # blocks until all scheduled jobs are done
OS-level scheduling (recommended for real automation)¶
For anything that should survive reboots, let the OS run it:
- Windows — Task Scheduler, or from PowerShell:
- Linux/macOS — cron. Run
crontab -eand add: (This runs the script every day at 02:00.)
Which to choose: use OS scheduling for periodic maintenance jobs (backups, reports). Use in-process scheduling only when the timing logic is part of a larger running application.
7. Turn a Script into a CLI Tool¶
Add arguments and help so your script is reusable by others.
argparse (standard library) turns a script into a proper command-line tool with --flags, help text, and validation.
import argparse
def main() -> None:
parser = argparse.ArgumentParser(description="Organize files by extension.")
parser.add_argument("folder", help="folder to organize")
parser.add_argument("--dry-run", action="store_true",
help="show what would happen without moving files")
parser.add_argument("--count", type=int, default=1,
help="number of times to repeat")
args = parser.parse_args()
print(f"folder={args.folder}, dry_run={args.dry_run}, count={args.count}")
if __name__ == "__main__":
main()
Now the script has a real interface:
Argument types:
- Positional (
"folder") — required, order matters. - Optional flag (
--dry-runwithaction="store_true") —Trueif present,Falseotherwise. - Typed option (
--countwithtype=int) — argparse validates and converts for you.
argparse vs Click vs Typer
argparse needs no install and covers most needs. For bigger tools with subcommands, look at Click or Typer (covered in the CLI Tools topic) — they reduce boilerplate and add niceties like colored output.
8. Logging Your Automation¶
Record what ran and what failed so unattended jobs are debuggable.
Unattended scripts (scheduled jobs) need logs — you won't be watching the terminal when they run. Use logging, not print.
import logging
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
handlers=[
logging.FileHandler("automation.log", encoding="utf-8"),
logging.StreamHandler(), # also print to console
],
)
log = logging.getLogger(__name__)
log.info("Backup started")
log.warning("Disk usage above 80%%")
log.error("Failed to connect to server")
Why logging beats print:
- Levels — filter noise (
DEBUG,INFO,WARNING,ERROR) without deleting code. - Timestamps — automatic, so you know when something happened.
- Destinations — send to a file, the console, or both at once.
- Always on — a scheduled job's
printoutput usually vanishes; a log file persists.
Putting it together¶
A realistic maintenance script that combines file cleanup, backup, logging, and a CLI — showing how the individual pieces above fit into one unattended job.
A realistic maintenance script combining several pieces — clean old files, back up, and log:
import argparse
import logging
import shutil
import time
from datetime import datetime
from pathlib import Path
log = logging.getLogger("maintenance")
def clean_old_files(folder: Path, max_age_days: int) -> int:
"""Delete files older than `max_age_days`. Returns count deleted."""
cutoff = time.time() - max_age_days * 86400
deleted = 0
for f in folder.iterdir():
if f.is_file() and f.stat().st_mtime < cutoff:
f.unlink()
deleted += 1
log.info("Deleted old file: %s", f.name)
return deleted
def backup(source: Path, dest_root: Path) -> str:
stamp = datetime.now().strftime("%Y%m%d-%H%M%S")
archive = shutil.make_archive(str(dest_root / f"bak-{stamp}"), "zip", source)
log.info("Created backup: %s", archive)
return archive
def main() -> None:
logging.basicConfig(level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s")
parser = argparse.ArgumentParser(description="Daily maintenance job.")
parser.add_argument("folder")
parser.add_argument("--max-age", type=int, default=30)
args = parser.parse_args()
folder = Path(args.folder)
log.info("Maintenance started on %s", folder)
removed = clean_old_files(folder, args.max_age)
backup(folder, folder.parent)
log.info("Done. Removed %d old files.", removed)
if __name__ == "__main__":
main()
Quick reference¶
A lookup table mapping each common automation task to the standard-library tool and a one-line example — scan it when you know what you want to do but not which function does it.
| Task | Tool | Example |
|---|---|---|
| Join paths | pathlib.Path | Path("a") / "b" |
| Read/write text | Path.read_text / .write_text | p.write_text("x") |
| Find files | Path.glob / .rglob | base.rglob("*.py") |
| Copy tree | shutil.copytree | copytree(src, dst) |
| Zip a folder | shutil.make_archive | make_archive("b", "zip", "src") |
| Read CSV | csv.DictReader | for row in DictReader(fh) |
| Read/write JSON | json.loads / .dumps | json.dumps(d, indent=2) |
| Run a program | subprocess.run | run([...], check=True) |
| Read env var | os.environ.get | os.environ.get("KEY", "default") |
| CLI arguments | argparse | parser.add_argument(...) |
| Logging | logging | log.info("...") |
Practice exercises¶
- Write a script that finds every
.logfile under a folder (recursively) and reports the total size in megabytes. - Build a CLI tool with
argparsethat takes a folder and an extension, and prints how many matching files exist. Add a--deleteflag that removes them. - Write a backup function that keeps only the 5 most recent zip backups in a folder, deleting older ones.
- Use
subprocessto runpython --versionand parse out just the version number. - Combine
csvandlogging: read a CSV of tasks, process each row, and log a success or failure line per row.
💬 Discussion
Have a question about this topic? Found an error? Share your thoughts below.