words = [w.strip("'").lower() for w in words if w.strip("'")]
The function reads a text file, extracts words using regex (sequences of [a-zA-Z']+), strips leading/trailing apostrophes, lowercases, counts with collections.Counter, and returns results sorted by frequency (descending) then alphabetically.
word_freq.py:
#!/usr/bin/env python3
"""Count word frequency in a text file using collections.Counter."""
import sys
from collections import Counter
import re
def count_word_frequency(file_path: str, encoding: str = "utf-8") -> list[tuple[str, int]]:
"""
Read a text file, count word frequencies using collections.Counter,
and return results sorted by frequency (descending), then alphabetically.
"""
try:
with open(file_path, "r", encoding=encoding) as f:
text = f.read()
except FileNotFoundError:
print(f"Error: File '{file_path}' not found.", file=sys.stderr)
sys.exit(1)
except IOError as e:
print(f"Error reading file '{file_path}': {e}", file=sys.stderr)
sys.exit(1)
# Extract words using regex, strip surrounding apostrophes, lowercase
words = re.findall(r"[a-zA-Z']+", text)
words = [w.strip("'").lower() for w in words if w.strip("'")]
counter = Counter(words)
# Sort by frequency descending, then alphabetically ascending
sorted_words = sorted(counter.items(), key=lambda x: (-x[1], x[0]))
return sorted_words
def main() -> None:
if len(sys.argv) != 2:
print(f"Usage: {sys.argv[0]} <file>", file=sys.stderr)
sys.exit(1)
result = count_word_frequency(sys.argv[1])
for word, freq in result:
print(f"{word}: {freq}")
if __name__ == "__main__":
main()
Usage:
python3 word_freq.py <file>
### Test Suite (15 tests, all passing) | Test | What It Verifies | |------|-----------------| | `test_basic` | Simple word counting works | | `test_case_insensitivity` | `Hello` / `HELLO` → `hello` (all lowercased) | | `test_sort_order` | Sorted by freq desc, then alpha asc | | `test_punctuation_stripped` | `hello!!!`, `"python"` → clean words | | `test_contractions_preserved` | `don't`, `it's`, `let's` kept intact | | `test_quoted_words` | `'hello'` → `hello` (no leading/trailing `'`) | | `test_numbers_ignored` | Pure numbers yield empty result | | `test_mixed_numbers_and_words` | `hello123` → `hello` | | `test_empty_file` | Empty file → `[]` | | `test_file_not_found` | Missing file → `SystemExit` | | `test_hyphenated_words` | `well-known` → `well` + `known` | | `test_multiple_lines` | Multi-line files work | | `test_large_file` | 1000×"the", 500×"and", 250×"a" sorted correctly | | `test_special_characters` | `@#$%^&*()` → `[]` | | `test_unicode` | Non-ASCII chars don't crash | ### Edge Case Output Input (`test_edge_cases.txt`): ``` Hello!!! World?? "Python" isn't it? It's... great. Don't stop! Let's go. Empty line above. Case: A a A a B b B b C c Numbers123 ignored456 too. Hyphenated-word? That's one word. What about 'single quotes' and "double quotes"? ``` Output (correct): ``` a: 4 b: 4 c: 2 quotes: 2 word: 2 about: 1 above: 1 ... don't: 1 ... isn't: 1 it's: 1 let's: 1 ... that's: 1 ... ``` All 15 tests pass: `15 passed in 0.01s` ---
{"model": "gpt-4o", "problem_class": "python-count-word-freq", "result": "passed", "tests": 15}