Compute Text Similarity

Problem Write code that computes a similarity score between two pieces of text and justify the metric chosen.

Input / Output

  • Input: strings a and b.
  • Output: similarity in [0, 1] (or a ranked interpretation).

Constraints

  • Clarify the use case first — near-duplicate detection, semantic similarity, or fuzzy matching — the metric follows from it.

Example

  • "the hotel was clean" vs "a clean hotel" → high Jaccard/cosine overlap after tokenization; "cheap flight" vs "inexpensive airfare" → near-zero lexical overlap, needs embeddings for a high score.
asked …
LeaderboardSalaryAccount