Cherche (search in French) allows you to create a neural search pipeline using retrievers and pre-trained language models as rankers.

Raphael Sourty

Last update: Nov 29, 2022

Related tags

Text Data & NLP search retrieval reader neural-networks searching bm25 flashtext neural-search

Overview

Cherche

Neural search

Cherche (search in French) allows you to create a neural search pipeline using retrievers and pre-trained language models as rankers. Cherche is meant to be used with small to medium sized corpora. Cherche's main strength is its ability to build diverse and end-to-end pipelines.

Installation 🤖

pip install cherche

To install the development version:

pip install git+https://github.com/raphaelsty/cherche

Documentation 📜

Documentation is available here. It provides details about retrievers, rankers, pipelines, question answering, summarization, and examples.

QuickStart 💨

Documents 📑

Cherche allows findings the right document within a list of objects. Here is an example of a corpus.

from cherche import data

documents = data.load_towns()

documents[:3]
[{'id': 0,
  'title': 'Paris',
  'url': 'https://en.wikipedia.org/wiki/Paris',
  'article': 'Paris is the capital and most populous city of France.'},
 {'id': 1,
  'title': 'Paris',
  'url': 'https://en.wikipedia.org/wiki/Paris',
  'article': "Since the 17th century, Paris has been one of Europe's major centres of science, and arts."},
 {'id': 2,
  'title': 'Paris',
  'url': 'https://en.wikipedia.org/wiki/Paris',
  'article': 'The City of Paris is the centre and seat of government of the region and province of Île-de-France.'
  }]

Retriever ranker 🔍

Here is an example of a neural search pipeline composed of a TfIdf that quickly retrieves documents, followed by a ranking model. The ranking model sorts the documents produced by the retriever based on the semantic similarity between the query and the documents.

from cherche import data, retrieve, rank
from sentence_transformers import SentenceTransformer

# List of dicts
documents = data.load_towns()

# Retrieve on fields title and article
retriever = retrieve.TfIdf(key="id", on=["title", "article"], documents=documents, k=30)

# Rank on fields title and article
ranker = rank.Encoder(
    key = "id",
    on = ["title", "article"],
    encoder = SentenceTransformer("sentence-transformers/all-mpnet-base-v2").encode,
    k = 3,
    path = "encoder.pkl"
)

# Pipeline creation
search = retriever + ranker

search.add(documents=documents)

search("Bordeaux")
[{'id': 57, 'similarity': 0.69513476},
 {'id': 63, 'similarity': 0.6214991},
 {'id': 65, 'similarity': 0.61809057}]

Map the index to the documents to access their contents.

search += documents
search("Bordeaux")
[{'id': 57,
  'title': 'Bordeaux',
  'url': 'https://en.wikipedia.org/wiki/Bordeaux',
  'article': 'Bordeaux ( bor-DOH, French: [bɔʁdo] (listen); Gascon Occitan: Bordèu [buɾˈðɛw]) is a port city on the river Garonne in the Gironde department, Southwestern France.',
  'similarity': 0.69513476},
 {'id': 63,
  'title': 'Bordeaux',
  'url': 'https://en.wikipedia.org/wiki/Bordeaux',
  'article': 'The term "Bordelais" may also refer to the city and its surrounding region.',
  'similarity': 0.6214991},
 {'id': 65,
  'title': 'Bordeaux',
  'url': 'https://en.wikipedia.org/wiki/Bordeaux',
  'article': "Bordeaux is a world capital of wine, with its castles and vineyards of the Bordeaux region that stand on the hillsides of the Gironde and is home to the world's main wine fair, Vinexpo.",
  'similarity': 0.61809057}]

Retrieve 👻

Cherche provides different retrievers that filter input documents based on a query.

retrieve.Elastic
retrieve.TfIdf
retrieve.Lunr
retrieve.BM25Okapi
retrieve.BM25L
retrieve.Flash
retrieve.Encoder

Rank 🤗

Cherche rankers are compatible with SentenceTransformers models, Hugging Face sentence similarity models, Hugging Face zero shot classification models, and of course with your own models.

Summarization and question answering

Cherche provides modules dedicated to summarization and question answering. These modules are compatible with Hugging Face's pre-trained models and can be fully integrated into neural search pipelines.

Acknowledgements 👏

The BM25 models available in Cherche are wrappers around rank_bm25. Elastic retriever is a wrapper around Python Elasticsearch Client. TfIdf retriever is a wrapper around scikit-learn's TfidfVectorizer. Lunr retriever is a wrapper around Lunr.py. Flash retriever is a wrapper around FlashText. DPR and Encode rankers are wrappers dedicated to the use of the pre-trained models of SentenceTransformers in a neural search pipeline. ZeroShot ranker is a wrapper dedicated to the use of the zero-shot sequence classifiers of Hugging Face in a neural search pipeline.

Dev Team 💾

The Cherche dev team is made up of Raphaël Sourty and François-Paul Servant 🥳

Comments

Added spelling corrector object

Hello ! I added a spelling corrector base class as well as the original implementation of the Norvig spelling corrector. The spelling corrector can be fitted directly on the pipeline's documents with the '.add(documents)' method. I also provided an optional (defaults to False) external dictionary, the one originally used by Norvig.

I have no issue updating my code for improvements, so feel free to suggest any modification !

opened by NicolasBizzozzero 4
0.0.5
Pull request for Cherche version 0.0.5

RAG: add RAG generator for open domain question answering

RapidFuzzy: New blazzing fast retriever

Retrievers: Provide similarities for each retriever

Union & Intersection: Keep similarity scores
opened by raphaelsty 1
Batch processing
Retrieving documents with batch of queries can significantly speed up things. It is now available for few models using the development version via the batch method.

Models involved are:

TfIdf retriever

Encoder retriever (milvus + faiss)

Encoder ranker (milvus)

DPR retriever (milvus + faiss)

DPR ranker (milvus)

Recommend retriever

Batch is not yet compatible with pipelines.
enhancement
opened by raphaelsty 0
Cherche 1.0.0
Here is an essential update for Cherche. The update retains the previous API and is compatible with previous versions. 🥳

Main additions:

Added compatibility with two new open-source retrievers: Meilisearch and TypeSense.

Compatibility with the Milvus index to use the retriever.Encoder and retriever.DPR models on massive corpora.

Compatibility with the Milvus index to store ranker embeddings in a database rather than in memory.

Progress bar when pre-computing embeddings by Encoder, DPR retrievers and Encoder, DPR rankers.

All pipelines (voting, intersection, concatenation) produce a similarity score. To do so, the pipeline object applies a softmax to normalize the scores, thus allowing us to "compare" the scores of two distinct models.

Integration of collaborative filtering models via adding a Recommend retriever and a Recommend ranker (indexation via Faiss and compatible with Milvus) to consider users' preferences in the search.
opened by raphaelsty 0
"IndexError: index out of range in self "While adding documents to cherche pipeline

I'm using a cherche pipline built of a tfidf retriever with a sentencetransformer ranker as follows : search = (retriever + ranker) While trying to add documents to the pipeline (search.add(documents=documents), I got this error :

"""/usr/local/lib/python3.7/dist-packages/torch/nn/functional.py in embedding(input, weight, padding_idx, max_norm, norm_type, scale_grad_by_freq, sparse) 2181 # remove once script supports set_grad_enabled 2182 no_grad_embedding_renorm(weight, input, max_norm, norm_type) -> 2183 return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse) 2184 2185

IndexError: index out of range in self"""

opened by delmetni 0
incomplete doc about metrics

https://raphaelsty.github.io/cherche/examples/eval_pipeline/ you say that one can find the explnaation about metrics here https://amitness.com/2020/08/information-retrieval-evaluation/ but it doesn't say what "precision" and "r-precision" are.

opened by fpservant 0

Releases(1.0.1)

1.0.1(Oct 27, 2022)

Removed the dependency with grpcio that can cause problems during installation.
Source code(tar.gz)
Source code(zip)
1.0.0(Oct 26, 2022)
What's Changed

Here is an essential update for Cherche! 🥳

Added compatibility with two new open-source retrievers: Meilisearch and TypeSense.

Compatibility with the Milvus index to use the retriever.Encoder and retriever.DPR models on massive corpora.

Compatibility with the Milvus index to store ranker embeddings in a database rather than in memory.

Progress bar when pre-computing embeddings by Encoder, DPR retrievers and Encoder, DPR rankers.

The path parameter is no longer used.

All pipelines (voting, intersection, concatenation) produce a similarity score. To do so, the pipeline object applies a softmax to normalize the scores, thus allowing us to "compare" the scores of two distinct models.

Integration of collaborative filtering models via adding a Recommend retriever and a Recommend ranker (indexation via Faiss and compatible with Milvus) to consider users' preferences in the search.

Cherche is now fully compatible with large-scale corpora and deeply integrates collaborative filtering. Updates retains the previous API and is compatible with previous versions.
Source code(tar.gz)
Source code(zip)
0.1.0(Jun 16, 2022)

Added compatibility with the ONNX environment and quantization to significantly speed up sentence transformers and question answering models. 🏎

It is now possible to choose the type of index for the Encoder and DPR retrievers in order to process the largest corpora while using the GPU.
Source code(tar.gz)
Source code(zip)
0.0.9(Apr 13, 2022)

Voting operator dedicated to retrievers and rankers.
Source code(tar.gz)
Source code(zip)
0.0.8(Mar 7, 2022)

Avoid checking similarities in TF-IDF retrievers while filtering documents.
Source code(tar.gz)
Source code(zip)
0.0.7(Mar 7, 2022)
Significant improvement in the speed of the TF-IDF retriever using sparse CSC matrix.

The setup.py file loads the readme file as UTF-8.

Source code(tar.gz)
Source code(zip)
0.0.6(Mar 3, 2022)
Update documentation

Update retriever Encoder and DPR, path is optionnal

Add deployment documentation

Update similarity type

Avoid round similarity

Source code(tar.gz)
Source code(zip)
0.0.5(Feb 8, 2022)
Loading and Saving tutorial

Fuzzy retriever

Similarities everywhere (retrievers, union, intersection provide similarity scores)

RAG generation

Source code(tar.gz)
Source code(zip)
0.0.4(Jan 20, 2022)

Update of the encoder retriever and the DPR retriever. Documents in the Faiss index will not be duplicated. Query embeddings can now be pre-computed for ranker Encoder and ranker DPR to speed up evaluation without having to compute it again.
Source code(tar.gz)
Source code(zip)
0.0.3(Jan 13, 2022)
Adding:

Translation

DPR retriever

Source code(tar.gz)
Source code(zip)
0.0.2(Jan 12, 2022)

Update of the Cherche dependencies. The previous dependencies were too strict and restrictive as they were limited to a specific version for each package.
Source code(tar.gz)
Source code(zip)
0.0.1(Jan 8, 2022)

Source code(tar.gz)
Source code(zip)

Owner

Raphael Sourty

PhD Student @ IRIT and Renault

GitHub

Guide to using pre-trained large language models of source code

Large Models of Source Code I occasionally train and publicly release large neural language models on programs, including PolyCoder. Here, I describe

947 Dec 28, 2022

Implementation of Natural Language Code Search in the project CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

CodeBERT-Implementation In this repo we have replicated the paper CodeBERT: A Pre-Trained Model for Programming and Natural Languages. We are interest

4 Jul 1, 2022

BPEmb is a collection of pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE) and trained on Wikipedia.

BPEmb is a collection of pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE) and trained on Wikipedia. Its intended use is as input for neural models in natural language processing.

1.1k Jan 3, 2023

RoNER is a Named Entity Recognition model based on a pre-trained BERT transformer model trained on RONECv2

RoNER RoNER is a Named Entity Recognition model based on a pre-trained BERT transformer model trained on RONECv2. It is meant to be an easy to use, hi

9 Nov 7, 2022

PyTorch Implementation of "Bridging Pre-trained Language Models and Hand-crafted Features for Unsupervised POS Tagging" (Findings of ACL 2022)

Feature_CRF_AE Feature_CRF_AE provides a implementation of Bridging Pre-trained Language Models and Hand-crafted Features for Unsupervised POS Tagging

6 Apr 29, 2022

Must-read papers on improving efficiency for pre-trained language models.

89 Jan 3, 2023

The repository for the paper: Multilingual Translation via Grafting Pre-trained Language Models

Graformer The repository for the paper: Multilingual Translation via Grafting Pre-trained Language Models Graformer (also named BridgeTransformer in t

22 Dec 14, 2022

Prompt-learning is the latest paradigm to adapt pre-trained language models (PLMs) to downstream NLP tasks

Prompt-learning is the latest paradigm to adapt pre-trained language models (PLMs) to downstream NLP tasks, which modifies the input text with a textual template and directly uses PLMs to conduct pre-trained tasks. This library provides a standard, flexible and extensible framework to deploy the prompt-learning pipeline. OpenPrompt supports loading PLMs directly from huggingface transformers. In the future, we will also support PLMs implemented by other libraries.

2.3k Jan 8, 2023

Chinese Pre-Trained Language Models (CPM-LM) Version-I

CPM-Generate 为了促进中文自然语言处理研究的发展，本项目提供了 CPM-LM (2.6B) 模型的文本生成代码，可用于文本生成的本地测试，并以此为基础进一步研究零次学习/少次学习等场景。[项目首页] [模型下载] [技术报告] 若您想使用CPM-1进行推理，我们建议使用高效推理工具BMI

1.4k Jan 3, 2023

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

3.2k Dec 31, 2022

Code associated with the "Data Augmentation using Pre-trained Transformer Models" paper

Data Augmentation using Pre-trained Transformer Models Code associated with the Data Augmentation using Pre-trained Transformer Models paper Code cont

44 Dec 31, 2022

Coreference resolution for English, French, German and Polish, optimised for limited training data and easily extensible for further languages

Coreferee Author: Richard Paul Hudson, Explosion AI 1. Introduction 1.1 The basic idea 1.2 Getting started 1.2.1 English 1.2.2 French 1.2.3 German 1.2

70 Dec 12, 2022

DziriBERT: a Pre-trained Language Model for the Algerian Dialect

DziriBERT is the first Transformer-based Language Model that has been pre-trained specifically for the Algerian Dialect.

117 Jan 7, 2023

⚖️ A Statutory Article Retrieval Dataset in French.

A Statutory Article Retrieval Dataset in French This repository contains the Belgian Statutory Article Retrieval Dataset (BSARD), as well as the code

19 Nov 17, 2022

:mag: Transformers at scale for question answering & neural search. Using NLP via a modular Retriever-Reader-Pipeline. Supporting DPR, Elasticsearch, HuggingFace's Modelhub...

Haystack is an end-to-end framework for Question Answering & Neural search that enables you to ... ... ask questions in natural language and find gran

6.4k Jan 9, 2023

One Stop Anomaly Shop: Anomaly detection using two-phase approach: (a) pre-labeling using statistics, Natural Language Processing and static rules; (b) anomaly scoring using supervised and unsupervised machine learning.

One Stop Anomaly Shop (OSAS) Quick start guide Step 1: Get/build the docker image Option 1: Use precompiled image (might not reflect latest changes):

148 Dec 26, 2022

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset.

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset. TunBERT was applied to three NLP downstream tasks: Sentiment Analysis (SA), Tunisian Dialect Identification (TDI) and Reading Comprehension Question-Answering (RCQA)

72 Dec 9, 2022

A Neural Language Style Transfer framework to transfer natural language text smoothly between fine-grained language styles like formal/casual, active/passive, and many more. Created by Prithiviraj Damodaran. Open to pull requests and other forms of collaboration.

Styleformer A Neural Language Style Transfer framework to transfer natural language text smoothly between fine-grained language styles like formal/cas

431 Dec 19, 2022

Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipgrams (i.e patterns with one or more gaps, either of fixed or dynamic size) in a quick and memory-efficient way. At the core is the tool ``colibri-patternmodeller`` whi ch allows you to build, view, manipulate and query pattern models.

Colibri Core by Maarten van Gompel, [email protected], Radboud University Nijmegen Licensed under GPLv3 (See http://www.gnu.org/licenses/gpl-3.0.html

122 Nov 17, 2022

Cherche (search in French) allows you to create a neural search pipeline using retrievers and pre-trained language models as rankers.

Related tags

Overview

Cherche

Installation 🤖

Documentation 📜

QuickStart 💨

Documents 📑

Retriever ranker 🔍

Retrieve 👻

Rank 🤗

Summarization and question answering

Acknowledgements 👏

See also 👀

Dev Team 💾

Comments

Added spelling corrector object

0.0.5

Batch processing

Cherche 1.0.0

"IndexError: index out of range in self "While adding documents to cherche pipeline

incomplete doc about metrics

Releases(1.0.1)

1.0.1(Oct 27, 2022)

1.0.0(Oct 26, 2022)

What's Changed

Here is an essential update for Cherche! 🥳

0.1.0(Jun 16, 2022)

0.0.9(Apr 13, 2022)

0.0.8(Mar 7, 2022)

0.0.7(Mar 7, 2022)

0.0.6(Mar 3, 2022)

0.0.5(Feb 8, 2022)

0.0.4(Jan 20, 2022)

0.0.3(Jan 13, 2022)

0.0.2(Jan 12, 2022)

0.0.1(Jan 8, 2022)

Owner

Raphael Sourty

Guide to using pre-trained large language models of source code

Implementation of Natural Language Code Search in the project CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

BPEmb is a collection of pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE) and trained on Wikipedia.

RoNER is a Named Entity Recognition model based on a pre-trained BERT transformer model trained on RONECv2

PyTorch Implementation of "Bridging Pre-trained Language Models and Hand-crafted Features for Unsupervised POS Tagging" (Findings of ACL 2022)

Must-read papers on improving efficiency for pre-trained language models.

The repository for the paper: Multilingual Translation via Grafting Pre-trained Language Models

Prompt-learning is the latest paradigm to adapt pre-trained language models (PLMs) to downstream NLP tasks

Chinese Pre-Trained Language Models (CPM-LM) Version-I

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

Code associated with the "Data Augmentation using Pre-trained Transformer Models" paper

Coreference resolution for English, French, German and Polish, optimised for limited training data and easily extensible for further languages

DziriBERT: a Pre-trained Language Model for the Algerian Dialect

⚖️ A Statutory Article Retrieval Dataset in French.

:mag: Transformers at scale for question answering & neural search. Using NLP via a modular Retriever-Reader-Pipeline. Supporting DPR, Elasticsearch, HuggingFace's Modelhub...

One Stop Anomaly Shop: Anomaly detection using two-phase approach: (a) pre-labeling using statistics, Natural Language Processing and static rules; (b) anomaly scoring using supervised and unsupervised machine learning.

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset.

A Neural Language Style Transfer framework to transfer natural language text smoothly between fine-grained language styles like formal/casual, active/passive, and many more. Created by Prithiviraj Damodaran. Open to pull requests and other forms of collaboration.