A warehouse audit wants to find tables that are probably copies or forks of each other. Write similar_tables(schemas, threshold).
schemas maps a table name to the list of its column names. Lists can contain the same column twice and can mix case or spaces, so compare columns after removing surrounding spaces and lower-casing, and treat each table's columns as a set.
The similarity of two tables is the number of columns they share divided by the number of distinct columns either of them has. A table with no columns is ignored entirely.
Return every pair whose similarity is at least threshold (compare the exact similarity, not a rounded one) as a tuple (table_a, table_b, score): each pair once, with table_a < table_b, and score rounded to 3 decimal places. Sort by score descending, then table_a, then table_b.
Python 3.13 in your browser — the standard library plus pandas and numpy; no pip installs.