Unicode & Text Processing Competent¶
When you'd use this
Strings, encoding, bytes, regex, text normalization and real-world text handling.
Handle real-world text safely: encodings, normalization, case-folding for comparisons, and formatting — essential for i18n, parsing, and data cleaning.
str vs bytes — the fundamental distinction¶
str is human-readable text (Unicode code points); bytes is raw binary. Knowing which you hold — and encoding/decoding at the boundary — prevents the most common text bugs.
# str — sequence of Unicode code points (text)
text = "Hello, 世界! 🐍"
print(type(text)) # <class 'str'>
print(len(text)) # 11 characters
# bytes — sequence of raw bytes (binary data)
data = b"Hello, ASCII only"
print(type(data)) # <class 'bytes'>
print(len(data)) # 17 bytes
# Encoding: str → bytes
encoded = text.encode("utf-8")
print(type(encoded)) # <class 'bytes'>
print(len(encoded)) # 18 bytes (Chinese chars = 3 bytes each, emoji = 4)
print(encoded) # b'Hello, \xe4\xb8\x96\xe7\x95\x8c! \xf0\x9f\x90\x8d'
# Decoding: bytes → str
decoded = encoded.decode("utf-8")
print(decoded == text) # True
Encoding schemes¶
The common ways text maps to bytes (ASCII, Latin-1, the UTF family) and when each applies — in practice, use UTF-8 everywhere unless a legacy system forces otherwise.
| Encoding | Bytes/char | Coverage | Use case |
|---|---|---|---|
| ASCII | 1 | English only (0-127) | Legacy systems |
| Latin-1 (ISO-8859-1) | 1 | Western European | Old web pages |
| UTF-8 | 1-4 | All Unicode | Default everywhere |
| UTF-16 | 2-4 | All Unicode | Windows internals |
| UTF-32 | 4 | All Unicode | Fixed-width processing |
# UTF-8 is variable-width
print("A".encode("utf-8")) # b'A' (1 byte)
print("é".encode("utf-8")) # b'\xc3\xa9' (2 bytes)
print("中".encode("utf-8")) # b'\xe4\xb8\xad' (3 bytes)
print("🐍".encode("utf-8")) # b'\xf0\x9f\x90\x8d' (4 bytes)
# Handling encoding errors
text = "Café ☕"
text.encode("ascii", errors="replace") # b'Caf? ?'
text.encode("ascii", errors="ignore") # b'Caf '
text.encode("ascii", errors="xmlcharrefreplace") # b'Café ☕'
Unicode code points and names¶
Convert between characters and their numeric code points (ord/chr), write them by escape or name, and inspect their category — useful for parsing, validation, and emoji/symbol handling.
# Every character has a code point (integer) and a name
print(ord("A")) # 65
print(ord("🐍")) # 128013
print(chr(65)) # A
print(chr(128013)) # 🐍
print(hex(ord("中"))) # 0x4e2d
# Unicode escape sequences
print("\u0041") # A (4-digit hex)
print("\U0001F40D") # 🐍 (8-digit hex for chars > 0xFFFF)
print("\N{SNAKE}") # 🐍 (by name)
# Get character name
import unicodedata
print(unicodedata.name("🐍")) # SNAKE
print(unicodedata.name("é")) # LATIN SMALL LETTER E WITH ACUTE
print(unicodedata.category("A")) # Lu (Letter, uppercase)
print(unicodedata.category("3")) # Nd (Number, decimal digit)
print(unicodedata.category("!")) # Po (Punctuation, other)
Text normalization¶
The same visible character can have multiple byte representations; normalizing (NFC/NFD/NFKC/NFKD) collapses them to one form so comparisons and searches behave — always normalize user input before comparing.
The same visual character can have different byte representations:
import unicodedata
# "é" can be stored two ways:
composed = "\u00e9" # single code point: é
decomposed = "e\u0301" # e + combining acute accent
print(composed == decomposed) # False!
print(len(composed)) # 1
print(len(decomposed)) # 2
# But they LOOK identical when printed!
# Normalize to compare
nfc = unicodedata.normalize("NFC", decomposed) # compose
nfd = unicodedata.normalize("NFD", composed) # decompose
print(nfc == composed) # True
print(nfd == decomposed) # True
# Always normalize before comparing user input!
def safe_compare(a, b):
return unicodedata.normalize("NFC", a) == unicodedata.normalize("NFC", b)
Normalization forms¶
| Form | Action | Use case |
|---|---|---|
| NFC | Compose (é = single char) | Default choice — most compact |
| NFD | Decompose (é = e + accent) | Text analysis, searching |
| NFKC | Compatibility compose | Search normalization |
| NFKD | Compatibility decompose | Stripping formatting |
# NFKC normalizes visual equivalents
print(unicodedata.normalize("NFKC", "fi")) # fi (ligature → two chars)
print(unicodedata.normalize("NFKC", "①")) # 1 (circled → plain)
print(unicodedata.normalize("NFKC", "Ⅳ")) # IV (Roman numeral → letters)
String methods — complete reference¶
The core str toolbox grouped by purpose — searching, transforming, splitting/joining, and testing — the methods you'll use in nearly every program that touches text.
Searching¶
Locate or count substrings — find/index for position, count for frequency, startswith/endswith for prefixes/suffixes — the building blocks of any text parsing.
s = "Hello, World! Hello, Python!"
s.find("Hello") # 0 (first occurrence, -1 if not found)
s.rfind("Hello") # 14 (last occurrence)
s.index("Hello") # 0 (like find, but raises ValueError)
s.count("Hello") # 2
s.startswith("Hello") # True
s.endswith(("!", "?", ".")) # True (accepts tuple)
"Python" in s # True (membership test)
Transforming¶
Reshape strings — trim whitespace, change case, pad/align, and replace substrings — the everyday cleanup you do on user input and before display.
s = " Hello, World! "
s.strip() # "Hello, World!"
s.lstrip() # "Hello, World! "
s.rstrip() # " Hello, World!"
s.strip(" !") # "Hello, World" — strip these chars
"hello".upper() # "HELLO"
"HELLO".lower() # "hello"
"hello world".title() # "Hello World"
"hello world".capitalize() # "Hello world"
"Hello".swapcase() # "hELLO"
"hello".center(20, "-") # "-------hello--------"
"hello".ljust(20) # "hello "
"hello".rjust(20) # " hello"
"42".zfill(8) # "00000042"
"hello world".replace("world", "Python") # "hello Python"
"hello world".replace("l", "L", 1) # "heLlo world" (max 1 replacement)
str.title() — what it's for
title() upper-cases the first letter of every word and lower-cases the rest — handy for display formatting: cleaning up user-entered names ("alice smith".title() → "Alice Smith"), normalizing city/place names, or formatting headings and labels for a UI.
Gotcha: it treats any non-letter as a word boundary, so apostrophes and hyphens break words oddly: "o'brien".title() → "O'Brien" and "it's".title() → "It'S". For human names with apostrophes, prefer .capitalize() per word or the string.capwords helper.
Splitting and joining¶
Turn a string into a list and back — split/splitlines to parse delimited data, join to assemble output, partition to break on the first separator.
"a,b,c".split(",") # ['a', 'b', 'c']
"a b c".split() # ['a', 'b', 'c'] (splits on any whitespace)
"a,b,c,d".split(",", 2) # ['a', 'b', 'c,d'] (max 2 splits)
"a\nb\nc".splitlines() # ['a', 'b', 'c']
",".join(["a", "b", "c"]) # "a,b,c"
" ".join(["Hello", "World"]) # "Hello World"
"\n".join(lines) # join with newlines
# Partition — split into 3 parts
"user@host.com".partition("@") # ('user', '@', 'host.com')
"no-at-sign".partition("@") # ('no-at-sign', '', '')
Testing¶
Check what a string contains — all digits, all letters, valid identifier — before converting or using it, so you validate input instead of catching exceptions later.
"123".isdigit() # True
"abc".isalpha() # True
"abc123".isalnum() # True
" ".isspace() # True
"Hello".istitle() # True
"HELLO".isupper() # True
"hello".islower() # True
"var_name".isidentifier() # True
"print".iskeyword() # False (use keyword.iskeyword())
f-strings — advanced formatting¶
Embed expressions directly in string literals with precise control over decimals, padding, alignment, and number formatting — the modern, readable way to build output.
name = "Alice"
score = 95.678
items = [1, 2, 3]
# Basic
f"Hello, {name}!" # "Hello, Alice!"
# Expressions
f"2 + 2 = {2 + 2}" # "2 + 2 = 4"
f"Items: {len(items)}" # "Items: 3"
# Format specifiers
f"{score:.2f}" # "95.68" (2 decimal places)
f"{1234567:,}" # "1,234,567" (thousands separator)
f"{0.75:.1%}" # "75.0%" (percentage)
f"{'hello':>20}" # " hello" (right-align)
f"{'hello':<20}" # "hello " (left-align)
f"{'hello':^20}" # " hello " (center)
f"{'hello':*^20}" # "*******hello********" (fill char)
f"{42:#x}" # "0x2a" (hex with prefix)
f"{42:#b}" # "0b101010" (binary)
f"{42:08d}" # "00000042" (zero-padded)
# Self-documenting (Python 3.8+)
x = 42
f"{x = }" # "x = 42"
f"{x * 2 = }" # "x * 2 = 84"
# Multiline
message = (
f"Name: {name}\n"
f"Score: {score:.1f}\n"
f"Grade: {'A' if score >= 90 else 'B'}"
)
# Nested f-strings
width = 20
f"{'hello':^{width}}" # " hello "
Real-world text processing patterns¶
Three tasks you'll hit in real projects — generating URL slugs, detecting a file's encoding, and stripping accents — each combining the primitives above.
Slug generation (URL-safe strings)¶
Turn an arbitrary title into a clean, lowercase, hyphenated URL slug — the exact transformation blogs and CMSs apply to post titles to build readable links.
import re
import unicodedata
def slugify(text):
"""Convert text to URL-safe slug."""
text = unicodedata.normalize("NFKD", text)
text = text.encode("ascii", "ignore").decode("ascii")
text = text.lower()
text = re.sub(r"[^\w\s-]", "", text)
text = re.sub(r"[-\s]+", "-", text).strip("-")
return text
print(slugify("Hello, World! Café ☕")) # "hello-world-caf"
print(slugify("Python 3.13: What's New")) # "python-313-whats-new"
Detecting encoding¶
Guess the encoding of an unknown text file before decoding it — the practical fix when you receive data files from other systems with no declared charset.
# pip install chardet
import chardet
with open("mystery.txt", "rb") as f:
raw = f.read()
detected = chardet.detect(raw)
print(detected) # {'encoding': 'utf-8', 'confidence': 0.99, ...}
text = raw.decode(detected["encoding"])
Stripping accents¶
Reduce accented letters to their plain ASCII base (café → cafe) by decomposing and dropping the combining marks — useful for search, sorting, and legacy-system compatibility.
import unicodedata
def strip_accents(text):
nfd = unicodedata.normalize("NFD", text)
return "".join(c for c in nfd if unicodedata.category(c) != "Mn")
print(strip_accents("Café résumé naïve")) # "Cafe resume naive"
Practice Exercises¶
- Write a function that detects the encoding of a file and converts it to UTF-8.
- Build a text normalizer that lowercases, strips accents, removes punctuation and collapses whitespace.
- Write a slug generator that handles Unicode input (Chinese, Arabic, emojis).
- Parse a log file with regex and extract structured data (timestamp, level, message).
- Implement a simple template engine that replaces
{{variable}}placeholders with values from a dict. - Count emoji in a string by checking Unicode categories.
💬 Discussion
Have a question about this topic? Found an error? Share your thoughts below.