A Python framework for conversational search

Castorini

Last update: Oct 23, 2022

Related tags

Deep Learning conversational-search

Overview

Chatty Goose

Multi-stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting

Installation

Make sure Java 11+ and Python 3.7+ are installed
Install the chatty-goose PyPI module

pip install chatty-goose

If you are using T5 or BERT, make sure to install PyTorch 1.4.0 - 1.7.1 using your specific platform instructions. Note that PyTorch 1.8 is currently incompatible due to the transformers version we currently use. Also make sure to install the corresponding torchtext version.
Download the English model for spaCy

python -m spacy download en_core_web_sm

Quickstart Guide

The following example shows how to initialize a searcher and build a ConversationalQueryRewriter agent from scratch using HQE and T5 as first-stage retrievers, and a BERT reranker. To see a working example agent, see chatty_goose/agents/chat.py.

First, load a searcher

from pyserini.search import SimpleSearcher

# Option 1: load a prebuilt index
searcher = SimpleSearcher.from_prebuilt_index("INDEX_NAME_HERE")
# Option 2: load a local Lucene index
searcher = SimpleSearcher("PATH_TO_INDEX")

searcher.set_bm25(0.82, 0.68)

Next, initialize one or more first-stage CQR retrievers

from chatty_goose.cqr import Hqe, Ntr
from chatty_goose.settings import HqeSettings, NtrSettings

hqe = Hqe(searcher, HqeSettings())
ntr = Ntr(NtrSettings())

Load a reranker

from chatty_goose.util import build_bert_reranker

reranker = build_bert_reranker()

Create a new RetrievalPipeline

from chatty_goose.pipeline import RetrievalPipeline

rp = RetrievalPipeline(searcher, [hqe, ntr], searcher_num_hits=50, reranker=reranker)

And we're done! Simply call rp.retrieve(query) to retrieve passages, or call rp.reset_history() to reset the conversational history of the retrievers.

Running Experiments

Clone the repo and all submodules (git submodule update --init --recursive)
Clone and build Anserini for evaluation tools
Install dependencies

pip install -r requirements.txt

Follow the instructions under docs/cqr_experiments.md to run experiments using HQE, T5, or fusion.

Example Agent

To run an interactive conversational search agent with ParlAI, simply run chat.py. By default, we use the CAsT 2019 pre-built Pyserini index, but it is possible to specify other indexes using the --from_prebuilt flag. See the file for other possible arguments:

python -m chatty_goose.agents.chat

Alternatively, run the agent using ParlAI's command line interface:

python -m parlai interactive --model chatty_goose.agents.chat:ChattyGooseAgent

We also provide instructions to deploy the agent to Facebook Messenger using ParlAI under examples/messenger.

Comments

Add baselines for CAsT 2020
Need someone help to add CAsT 2020 baseline results:

[ ] Naive: CQR without canonical responses

[ ] Canonical: CQR with canonical (manual) response

CQR methods: HQE /Ntr (T5)
enhancement help wanted
opened by justram 2
Running HQE and getting the reformulated queries
Dear authors,

I am trying to use your method in some of my work. For that, I need to get the reformulated queries (instead of only the generated ranked hits).

I am trying to run the HQE experiment as indicated using:

python -m experiments.run_retrieval \ --experiment hqe \ --hits 1000 \ --sparse_index cast2019 \ --qid_queries $input_query_json \ --output ./output/hqe_bm25

However, when I print the arguments passed inside the retrieval pipeline (L101 of retrieval_pipeline.py) I get as query the raw/original/last-turn query string, and as manual_context_buffer[turn_id] simply None. If I'm not mistaken, that means that running the specific experiment equals to no reformulation being done at all. Can you check/confirm this?

Digging more into the code, it seems to me that the queries I'd like to access are inside cqr_queries, but still, it seems to me that context should be empty/None in that case - probably resulting to no reformulation done at all.
opened by littlewine 1
Query rewriting fix

Thank you with the project.

The fix to below will be hits = rp.retrieve(query, manual_context_buffer[turn_id-1] if turn_id!=0 else None), to pass the last previous canonical response.

https://github.com/castorini/chatty-goose/blob/f9c21c8b7b6194d11d7aec5b4e218174cde98418/experiments/run_retrieval.py#L100

opened by xeniaqian94 1
Update based on Pyserini==0.14.0 and fix canonical response bug
Main change:

change --dense_index from temporary one to pyserini prebuilt index name

fix canonical response bug, which previously add current response to context

since now we have --dense_index, change option name --index to --sparse_index
opened by jacklin64 0
Add chatty goose support for dense retrieval and hybrid search for T5 and CQE

New features added: (only for T5 and CQE, may consider HQE in the future) (1) Dense retrieval (2) Dense-sparse hybrid retrieval

Some arg might be confused and may be changed in the future: (1) --index, --dense_index: may change to --sparse_index and --dense_index (2) --experiment now has options (hqe,cqe,t5,fusion,cqe_t5_fusion) may change to (hqe,cqe,t5,hqe_t5fusion,cqe_t5_fusion)

opened by jacklin64 0
Add cast2020 baseline

This PR adds both naive and canonical baselines for CAsT2020 topics. The results are overall lower as compared to CAst2019 and the results from the canonical run are only slightly better for some metrics as compared to results from the naive run.

Resolves #23

opened by saileshnankani 0
Add support for canonical response

This PR adds support for using manual_canonical_result_id in the CAsT2020 data for both ntr and hqe (for #23).

For ntr, rewrite uses the passage corresponding to the canonical document in the history. We only use 1 passage in the historical context as otherwise, it exceed 512 tokens limit. For e.g., it uses q1/P1/q2 and then q1/q2/P2/q3 and so on.
enhancement

opened by saileshnankani 0
CQR Replication
Add CQR replication for Fusion BM25

Library versions used: torch==1.7.0 torchvision==0.8.1 torchtext==0.8

Results:

map all 0.2584 recall_1000 all 0.8028 ndcg_cut_1 all 0.3353 ndcg_cut_3 all 0.3247

Details and reproduction results can be found in the notebook
opened by saileshnankani 0
Rename classes and update messenger bot
Breaking changes:

Renamed several classes to follow Python conventions / be more consistent

chatty_goose.agents.cqragent -> chatty_goose.agents.chat

HQE -> Hqe

T5_NTR -> Ntr

HQESettings -> HqeSettings

T5Settings -> NtrSettings

CQRType -> CqrType

CQRSettings -> CqrSettings

CQR -> ConversationalQueryRewriter
opened by edwinzhng 0

document spaCy model dependency

With a fresh install, we get the following error if we try to run anything:

OSError: [E050] Can't find model 'en_core_web_sm'. It doesn't seem to be a shortcut link, a Python package or a valid path to a data directory.

Solution is:

$ python -m spacy download en_core_web_sm

We should document this.

opened by lintool 0

PyTorch version: needs Torch 1.7 (won't work with 1.8)
With a from-scratch installation, the module pulls in Torch 1.8, which causes this error:

ImportError: cannot import name 'SAVE_STATE_WARNING' from 'torch.optim.lr_scheduler' (/anaconda3/envs/chatty-goose-test/lib/python3.7/site-packages/torch/optim/lr_scheduler.py)

Downgrading fixes the issue:

$ pip install torch==1.7.1 torchtext==0.8.1

Should we pin the version in our module dependencies? Or at the very least this needs to be documented.=
opened by lintool 0

dependency conflict

Hi,

When I install chatty-goose from github using:

python -m pip install git+https://github.com/castorini/chatty-goose.git

I met this issue:

ERROR: Cannot install chatty-goose and chatty-goose==0.2.0 because these package versions have conflicting dependencies.

The conflict is caused by:
    chatty-goose 0.2.0 depends on pyserini==0.14.0
    pygaggle 0.0.3.1 depends on pyserini==0.10.1.0

It seems that chatty-goose requires pyserini==0.14.0 as well as pygaggle 0.0.3.1. However, pygaggle 0.0.3.1 and pyserini==0.14.0 do not play nice with each other

Could someone provide some help?

Thanks!

opened by dayuyang1999 1

Expansion to new datasets
Does it make sense to expand Chatty Goose to new datasets? For example:

MANtIS - a multi-domain information seeking dialogues dataset: https://guzpenha.github.io/MANtIS/

ClariQ - Search-oriented Conversational AI (SCAI) EMNLP https://github.com/aliannejadi/ClariQ

enhancement
opened by lintool 1
Checkpoint transformation

According to @edwinzhng's replication log, we have a reranker checkpoint mismatch issue. Currently, we have diffs in our reranking model and the pygaggle's default model.

Related to this issue: I think we need a folder to put/track our tf2torch ckpt transformation/sanity check scripts?

opened by justram 0

Releases(0.2.0)

0.2.0(May 7, 2021)

Breaking changes

Renamed several classes to follow Python conventions / be more consistent

chatty_goose.agents.cqragent -> chatty_goose.agents.chat HQE -> Hqe T5_NTR -> Ntr HQESettings -> HqeSettings T5Settings -> NtrSettings CQRType -> CqrType CQRSettings -> CqrSettings CQR -> ConversationalQueryRewriter
Source code(tar.gz)
Source code(zip)
v0.1.0(Mar 8, 2021)
BREAKING CHANGES

Integrate ParlAI Facebook Messenger example for a demo by @edwinzhng

Integrate Pyserini/Pygaggle for a reference implementation of multi-stage passage retrieval by @edwinzhng

Add replication log for TREC CAsT 2019 conversational passage retrieval task by @edwinzhng

Source code(tar.gz)
Source code(zip)

Owner

Castorini

Deep learning for natural language processing and information retrieval at the University of Waterloo

GitHub

PyTorch implementation for ACL 2021 paper "Maria: A Visual Experience Powered Conversational Agent".

Maria: A Visual Experience Powered Conversational Agent This repository is the Pytorch implementation of our paper "Maria: A Visual Experience Powered

22 Dec 12, 2022

This is the repo for our work "Towards Persona-Based Empathetic Conversational Models" (EMNLP 2020)

Towards Persona-Based Empathetic Conversational Models (PEC) This is the repo for our work "Towards Persona-Based Empathetic Conversational Models" (E

35 Nov 17, 2022

The Adapter-Bot: All-In-One Controllable Conversational Model

The Adapter-Bot: All-In-One Controllable Conversational Model This is the implementation of the paper: The Adapter-Bot: All-In-One Controllable Conver

37 Nov 4, 2022

NUANCED is a user-centric conversational recommendation dataset that contains 5.1k annotated dialogues and 26k high-quality user turns.

NUANCED: Natural Utterance Annotation for Nuanced Conversation with Estimated Distributions Overview NUANCED is a user-centric conversational recommen

18 Dec 28, 2021

ICLR 2021: Pre-Training for Context Representation in Conversational Semantic Parsing

SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing This repository contains code for the ICLR 2021 paper "SCoRE: Pre-Tr

28 Oct 2, 2022

Facestar dataset. High quality audio-visual recordings of human conversational speech.

Facestar Dataset Description Existing audio-visual datasets for human speech are either captured in a clean, controlled environment but contain only a

87 Dec 21, 2022

Model search is a framework that implements AutoML algorithms for model architecture search at scale

Model search (MS) is a framework that implements AutoML algorithms for model architecture search at scale. It aims to help researchers speed up their exploration process for finding the right model architecture for their classification problems (i.e., DNNs with different types of layers).

3.2k Dec 31, 2022

⚡ Fast • 🪶 Lightweight • 0️⃣ Dependency • 🔌 Pluggable • 😈 TLS interception • 🔒 DNS-over-HTTPS • 🔥 Poor Man's VPN • ⏪ Reverse & ⏩ Forward • 👮🏿 "Proxy Server" framework • 🌐 "Web Server" framework • ➵ ➶ ➷ ➠ "PubSub" framework • 👷 "Work" acceptor & executor framework

Table of Contents Features Install Using PIP Stable version Development version Using Docker Stable version Development version Using HomeBrew Stable

2.2k Jan 8, 2023

TorchPQ is a python library for Approximate Nearest Neighbor Search (ANNS) and Maximum Inner Product Search (MIPS) on GPU using Product Quantization (PQ) algorithm.

Efficient implementations of Product Quantization and its variants using Pytorch and CUDA

146 Dec 28, 2022

This repo contains the code and data used in the paper "Wizard of Search Engine: Access to Information Through Conversations with Search Engines"

Wizard of Search Engine: Access to Information Through Conversations with Search Engines by Pengjie Ren, Zhongkun Liu, Xiaomeng Song, Hongtao Tian, Zh

19 Oct 27, 2022

Deep Text Search is an AI-powered multilingual text search and recommendation engine with state-of-the-art transformer-based multilingual text embedding (50+ languages).

Deep Text Search - AI Based Text Search & Recommendation System Deep Text Search is an AI-powered multilingual text search and recommendation engine w

19 Sep 29, 2022

Densely Connected Search Space for More Flexible Neural Architecture Search (CVPR2020)

DenseNAS The code of the CVPR2020 paper Densely Connected Search Space for More Flexible Neural Architecture Search. Neural architecture search (NAS)

291 Nov 18, 2022

The deployment framework aims to provide a simple, lightweight, fast integrated, pipelined deployment framework that ensures reliability, high concurrency and scalability of services.

savior是一个能够进行快速集成算法模块并支持高性能部署的轻量开发框架。能够帮助将团队进行快速想法验证（PoC），避免重复的去github上找模型然后复现模型；能够帮助团队将功能进行流程拆解，很方便的提高分布式执行效率；能够有效减少代码冗余，减少不必要负担。

125 Dec 22, 2022

FEDn is an open-source, modular and ML-framework agnostic framework for Federated Machine Learning

FEDn is an open-source, modular and ML-framework agnostic framework for Federated Machine Learning (FedML) developed and maintained by Scaleout Systems. FEDn enables highly scalable cross-silo and cross-device use-cases over FEDn networks.

75 Nov 9, 2022

Naszilla is a Python library for neural architecture search (NAS)

A repository to compare many popular NAS algorithms seamlessly across three popular benchmarks (NASBench 101, 201, and 301). You can implement your ow

270 Jan 3, 2023

Code for "Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search"

Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search This is an implementation for our paper Contextual Non-Loca

50 Dec 3, 2022

[ICLR 2021] "Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective" by Wuyang Chen, Xinyu Gong, Zhangyang Wang

Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective [PDF] Wuyang Chen, Xinyu Gong, Zhangyang Wang In ICLR 2

156 Nov 28, 2022

Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search

CLIP-GLaSS Repository for the paper Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search An in-browser demo is

172 Dec 22, 2022

Seach Losses of our paper 'Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search', accepted by ICLR 2021.

CSE-Autoloss Designing proper loss functions for vision tasks has been a long-standing research direction to advance the capability of existing models

54 Dec 17, 2022

A Python framework for conversational search

Related tags

Overview

Chatty Goose

Multi-stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting

Installation

Quickstart Guide

Running Experiments

Example Agent

Comments

Releases(0.2.0)

0.2.0(May 7, 2021)

Breaking changes

v0.1.0(Mar 8, 2021)

BREAKING CHANGES

Owner

Castorini

PyTorch implementation for ACL 2021 paper "Maria: A Visual Experience Powered Conversational Agent".

This is the repo for our work "Towards Persona-Based Empathetic Conversational Models" (EMNLP 2020)

The Adapter-Bot: All-In-One Controllable Conversational Model

NUANCED is a user-centric conversational recommendation dataset that contains 5.1k annotated dialogues and 26k high-quality user turns.

ICLR 2021: Pre-Training for Context Representation in Conversational Semantic Parsing

Facestar dataset. High quality audio-visual recordings of human conversational speech.

Model search is a framework that implements AutoML algorithms for model architecture search at scale

TorchPQ is a python library for Approximate Nearest Neighbor Search (ANNS) and Maximum Inner Product Search (MIPS) on GPU using Product Quantization (PQ) algorithm.

This repo contains the code and data used in the paper "Wizard of Search Engine: Access to Information Through Conversations with Search Engines"

Deep Text Search is an AI-powered multilingual text search and recommendation engine with state-of-the-art transformer-based multilingual text embedding (50+ languages).

Densely Connected Search Space for More Flexible Neural Architecture Search (CVPR2020)

The deployment framework aims to provide a simple, lightweight, fast integrated, pipelined deployment framework that ensures reliability, high concurrency and scalability of services.

FEDn is an open-source, modular and ML-framework agnostic framework for Federated Machine Learning

Naszilla is a Python library for neural architecture search (NAS)

Code for "Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search"

[ICLR 2021] "Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective" by Wuyang Chen, Xinyu Gong, Zhangyang Wang

Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search

Seach Losses of our paper 'Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search', accepted by ICLR 2021.