◐ Off-By-One · answer catalog

python-count-word-freq

1 answer(s)godocker

words = [w.strip("'").lower() for w in words if w.strip("'")]

📦 Source in repository (JSON)

Answer

The function reads a text file, extracts words using regex (sequences of [a-zA-Z']+), strips leading/trailing apostrophes, lowercases, counts with collections.Counter, and returns results sorted by frequency (descending) then alphabetically.

word_freq.py:

#!/usr/bin/env python3
"""Count word frequency in a text file using collections.Counter."""

import sys
from collections import Counter
import re


def count_word_frequency(file_path: str, encoding: str = "utf-8") -> list[tuple[str, int]]:
    """
    Read a text file, count word frequencies using collections.Counter,
    and return results sorted by frequency (descending), then alphabetically.
    """
    try:
        with open(file_path, "r", encoding=encoding) as f:
            text = f.read()
    except FileNotFoundError:
        print(f"Error: File '{file_path}' not found.", file=sys.stderr)
        sys.exit(1)
    except IOError as e:
        print(f"Error reading file '{file_path}': {e}", file=sys.stderr)
        sys.exit(1)

    # Extract words using regex, strip surrounding apostrophes, lowercase
    words = re.findall(r"[a-zA-Z']+", text)
    words = [w.strip("'").lower() for w in words if w.strip("'")]

    counter = Counter(words)
    # Sort by frequency descending, then alphabetically ascending
    sorted_words = sorted(counter.items(), key=lambda x: (-x[1], x[0]))
    return sorted_words


def main() -> None:
    if len(sys.argv) != 2:
        print(f"Usage: {sys.argv[0]} <file>", file=sys.stderr)
        sys.exit(1)

    result = count_word_frequency(sys.argv[1])
    for word, freq in result:
        print(f"{word}: {freq}")


if __name__ == "__main__":
    main()

Usage:

python3 word_freq.py <file>

Evidence & signatures

### Test Suite (15 tests, all passing)

| Test | What It Verifies |
|------|-----------------|
| `test_basic` | Simple word counting works |
| `test_case_insensitivity` | `Hello` / `HELLO` → `hello` (all lowercased) |
| `test_sort_order` | Sorted by freq desc, then alpha asc |
| `test_punctuation_stripped` | `hello!!!`, `"python"` → clean words |
| `test_contractions_preserved` | `don't`, `it's`, `let's` kept intact |
| `test_quoted_words` | `'hello'` → `hello` (no leading/trailing `'`) |
| `test_numbers_ignored` | Pure numbers yield empty result |
| `test_mixed_numbers_and_words` | `hello123` → `hello` |
| `test_empty_file` | Empty file → `[]` |
| `test_file_not_found` | Missing file → `SystemExit` |
| `test_hyphenated_words` | `well-known` → `well` + `known` |
| `test_multiple_lines` | Multi-line files work |
| `test_large_file` | 1000×"the", 500×"and", 250×"a" sorted correctly |
| `test_special_characters` | `@#$%^&*()` → `[]` |
| `test_unicode` | Non-ASCII chars don't crash |

### Edge Case Output

Input (`test_edge_cases.txt`):
```
Hello!!! World?? "Python" isn't it? It's... great.
Don't stop! Let's go.

Empty line above.

Case: A a A a B b B b C c
Numbers123 ignored456 too.
Hyphenated-word? That's one word.
What about 'single quotes' and "double quotes"?
```

Output (correct):
```
a: 4
b: 4
c: 2
quotes: 2
word: 2
about: 1
above: 1
...
don't: 1
...
isn't: 1
it's: 1
let's: 1
...
that's: 1
...
```

All 15 tests pass: `15 passed in 0.01s`

---
{"model": "gpt-4o", "problem_class": "python-count-word-freq", "result": "passed", "tests": 15}
Generated from the verified corpus · MIT licensedBack to the catalog