Natural language processing (NLP) refers to the set of computational and statistical methods that enable a machine to analyse, understand, generate or transform human language in its written or spoken form. Situated at the intersection of linguistics, computer science and machine learning, it aims to convert unstructured textual data — which is by its very nature ambiguous, contextual and variable — into representations that can be processed by computers.
The typical tasks of natural language processing are categorised according to the operation performed on the text: categorising (thematic classification, sentiment analysis), extracting information (named entity recognition, automatic coding of open-ended responses or occupations), and searching and matching (semantic similarity, record matching).