Phase 6 · Professional Database EngineeringModule 42~72 min read

Full-Text Search

Build language-aware search with tsvector, tsquery, ranking, highlighting, and GIN indexes.

What you'll learn

Build language-aware search with tsvector, tsquery, ranking, highlighting, and GIN indexes. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply Text search documents to a realistic data question
  • Apply tsvector to a realistic data question
  • Apply tsquery to a realistic data question
  • Apply Ranking to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
tsvectorNormalized lexemes with optional positions and weightsBuild from the exact searchable document and language configuration
tsqueryA normalized search expressionParse user intent with the appropriate safe query constructor
GIN indexAn inverted index over tokensUse for scalable containment and search while budgeting write cost

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the language-aware search question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Rank and highlight matching books

A generated search document can be indexed, then a web-style query is matched and ranked consistently.

book_search.sql
SELECT id, title,
  ts_rank(search_document, query) AS rank,
  ts_headline('english', description, query) AS excerpt
FROM catalog.books,
  websearch_to_tsquery('english', 'transaction isolation') AS query
WHERE search_document @@ query
ORDER BY rank DESC, id
LIMIT 20;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

ILIKE '%term%' is substring matching, not language-aware search, and a normal B-tree index cannot accelerate its leading wildcard.

Independent workshop

Build a review-ready language-aware search lab against the course commerce dataset.

Your finished workshop must include:

  • Text search documents
  • tsvector
  • tsquery
  • Ranking
  • Highlighting
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • tsvector: Build from the exact searchable document and language configuration
  • tsquery: Parse user intent with the appropriate safe query constructor
  • GIN index: Use for scalable containment and search while budgeting write cost

Quick check

1. Which rule best applies to tsvector?

2. Which rule best applies to tsquery?

3. Which rule best applies to GIN index?

Next: Table Partitioning & Data Lifecycle