GyaanamKnowledge for All
Back to Science & TechnologyAll concepts

Optical Character Recognition

Syllabusissues relating to intellectual property rights

Science & TechnologyPublished 6 August 2026

Optical Character Recognition (OCR) converts text visible in a scanned image into machine-readable, editable and searchable text. It detects printed or handwritten symbols represented by pixels and maps them to encoded characters, rather than merely storing their visual appearance.

Image preparation and segmentation

OCR first improves the scanned page so that text can be separated reliably from the background.

  • The page is converted to grayscale or black and white through binarisation, while noise, uneven illumination and distortions are reduced.
  • Skew correction aligns tilted lines, while layout analysis identifies columns, paragraphs, illustrations and reading order.
  • Segmentation divides text regions into lines, words or characters, although modern sequence models may recognize an entire line without separating every character.

Recognition and error correction

The recognition system compares visual patterns with learned representations of letters, numerals and punctuation.

  • During feature extraction, the system identifies useful patterns such as strokes, curves, intersections and spatial relationships.
  • A statistical or neural-network classifier assigns the most likely character or character sequence to each image region.
  • A dictionary or language model uses spelling and context to correct improbable sequences, while confidence scores flag uncertain results for review.

Output, limitations and legal significance

The recognized characters are stored in formats such as Unicode text or as a searchable text layer linked to the original page image.

  • Damaged pages, unusual fonts, ligatures, complex scripts, marginal notes, tables and multi-column layouts can reduce accuracy, making human proofreading necessary.
  • OCR recognizes character patterns, not the meaning or factual correctness of the text; errors can therefore affect search, quotation and downstream data use.
  • OCR changes the format of a work but not its copyright status. Reproduction and permitted exceptions remain governed by the Copyright Act, 1957, particularly Sections 14 and 52.

How UPSC asks this

Mains

UPSC may ask how OCR supports digitisation, accessibility and language technologies, while requiring discussion of accuracy, preservation and copyright concerns in converting books into machine-readable datasets.

Keep reading

The news behind topics like this, explained every morning

Every morning Gyaanam reads The Hindu, the Indian Express and PIB and picks what matters for UPSC. Each story is written up against the syllabus line it belongs to. Your first 15 days are free.

Sign up