Semantic Annotation

Semantic Annotation


Semantic annotation is marking up what the expressions inside a piece of text actually refer to, in a form a machine can read. What gets recorded is the meaning rather than the word. An annotation settles whether "Ford" is a brand, a surname or a place, and usually ties it to a unique identifier in a knowledge base.

The process runs in two steps. First the candidates in the text are found, which is the named entity recognition step. Then each candidate is linked to a record in an ontology or knowledge base. That is how "Ford" and "Ford Motor Company" arrive at the same identifier and the system knows two spellings mean one thing. On the web, schema.org markup does the same job by binding page fields such as price, author and event date to types a search engine understands.

It is needed wherever text has to be structured for search or inference. Enterprise search, product catalogues, content recommendation, legal and medical document processing and search visibility all qualify.

A concrete case: a news archive holds 12,000 stories mentioning "Apple". Without annotation a query returns both the company and the fruit. Once entities are linked to identifiers, the company stories group into their own set and the archive query returns something usable.

Annotation is not cheap. Done by hand it takes time, done automatically it still needs human review on the ambiguous cases.

From generative AI strategy to custom agent development and retrieval architectures, we help you scale AI responsibly.
Discuss your AI project