Kroomsa: A search engine for the curious

Wingify

Last update: Jun 20, 2022

Related tags

Overview

Kroomsa

A search engine for the curious. It is a search algorithm designed to engage users by exposing them to relevant yet interesting content during their session.

Description

The search algorithm implemented in your website greatly influences visitor engagement. A decent implementation can significantly reduce dependency on standard search engines like Google for every query thus, increasing engagement. Traditional methods look at terms or phrases in your query to find relevant content based on syntactic matching. Kroomsa uses semantic matching to find content relevant to your query. There is a blog post expanding upon Kroomsa's motivation and its technical aspects.

Getting Started

Prerequisites

Python 3.6.5
Run the project directory setup: python3 ./setup.py in the root directory.
Tensorflow's Universal Sentence Encoder 4
- The model is available at this link. Download the model and extract the zip file in the /vectorizer directory.
MongoDB is used as the database to collate Reddit's submissions. MongoDB can be installed following this link.
To fetch comments of the reddit submissions, PRAW is used. To scrape credentials are needed that authorize the script for the same. This is done by creating an app associated with a reddit account by following this link. For reference you can follow this tuorial written by Shantnu Tiwari.
- Register multiple instances and retrieve their credentials, then add them to the /config under bot_codes parameter in the following format: "client_id client_secret user_agent" as list elements separated by ,.
Docker-compose (For dockerized deployment only): Install the latest version following this link.

Installing

Create a python environment and install the required packages for preprocessing using: python3 -m pip install -r ./preprocess_requirements.txt
Collating a dataset of Reddit submissions
- Scraping posts
  - Pushshift's API is being used to fetch Reddit submissions. In the root directory, run the following command: python3 ./pre_processing/scraping/questions/scrape_questions.py. It launches a script that scrapes the subreddits sequentially till their inception and stores the submissions as JSON objects in /pre_processing/scraping/questions/scraped_questions. It then partitions the scraped submissions into as many equal parts as there are registered instances of bots.
- Scraping comments
  - After populating the configuration with bot_codes, we can begin scraping the comments using the partitioned submission files created while scraping submissions. Using the following command: python3 ./pre_processing/scraping/comments/scrape_comments.py multiple processes are spawned that fetch comment streams simultaneously.
- Insertion
  - To insert the submissions and associated comments, use the following commands: python3 ./pre_processing/db_insertion/insertion.py. It inserts the posts and associated comments in mongo.
  - To clean the comments and tag the posts that aren't public due to any reason, Run python3 ./post_processing/post_processing.py. Apart from cleaning, it also adds emojis to each submission object (This behavior is configurable).
Creating a FAISS Index
- To create a FAISS index, run the following command: python3 ./index/build_index.py. By default, it creates an exhaustive IDMap, Flat index but is configurable through the /config.
Database dump (For dockerized deployment)
- For dockerized deployment, a database dump is required in /mongo_dump. Use the following command at the root dir to create a database dump. mongodump --db database_name(default: red) --collection collection_name(default: questions) -o ./mongo_dump.

Execution

Local deployment (Using Gunicorn)
- Create a python environment and install the required packages using the following command: python3 -m pip install -r ./inference_requirements.txt
- A local instance of Kroomsa can be deployed using the following command: gunicorn -c ./gunicorn_config.py server:app
Dockerized demo
- Set the demo_mode to True in /config.
- Build images: docker-compose build
- Deploy: docker-compose up

Authors

License

This project is licensed under the Apache License Version 2.0

Apache Spark - A unified analytics engine for large-scale data processing

Apache Spark Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R, and an op

34.7k Jan 4, 2023

CPU inference engine that delivers unprecedented performance for sparse models

The DeepSparse Engine is a CPU runtime that delivers unprecedented performance by taking advantage of natural sparsity within neural networks to reduce compute required as well as accelerate memory bound workloads. It is focused on model deployment and scaling machine learning pipelines, fitting seamlessly into your existing deployments as an inference backend.

1.2k Jan 9, 2023

PPLNN is a Primitive Library for Neural Network is a high-performance deep-learning inference engine for efficient AI inferencing

943 Jan 7, 2023

.NET bindings for the Pytorch engine

TorchSharp TorchSharp is a .NET library that provides access to the library that powers PyTorch. It is a work in progress, but already provides a .NET

17 Aug 30, 2021

Brax is a differentiable physics engine that simulates environments made up of rigid bodies, joints, and actuators

Brax is a differentiable physics engine that simulates environments made up of rigid bodies, joints, and actuators. It's also a suite of learning algorithms to train agents to operate in these environments (PPO, SAC, evolutionary strategy, and direct trajectory optimization are implemented).

1.5k Jan 2, 2023

Kroomsa: A search engine for the curious

Related tags

Overview

Kroomsa

Description

Getting Started

Prerequisites

Installing

Execution

Authors

License

You might also like...

Apache Spark - A unified analytics engine for large-scale data processing

CPU inference engine that delivers unprecedented performance for sparse models

PPLNN is a Primitive Library for Neural Network is a high-performance deep-learning inference engine for efficient AI inferencing

.NET bindings for the Pytorch engine

Brax is a differentiable physics engine that simulates environments made up of rigid bodies, joints, and actuators

QSYM: A Practical Concolic Execution Engine Tailored for Hybrid Fuzzing

YoHa - A practical hand tracking engine.

On-device speech-to-intent engine powered by deep learning

On-device speech-to-index engine powered by deep learning.

Owner

Wingify

This is the pytorch code for the paper Curious Representation Learning for Embodied Intelligence.

This repo contains the code and data used in the paper "Wizard of Search Engine: Access to Information Through Conversations with Search Engines"

Deep Text Search is an AI-powered multilingual text search and recommendation engine with state-of-the-art transformer-based multilingual text embedding (50+ languages).

For AILAB: Cross Lingual Retrieval on Yelp Search Engine

A fast, dataset-agnostic, deep visual search engine for digital art history

Udacity's CS101: Intro to Computer Science - Building a Search Engine

Model search is a framework that implements AutoML algorithms for model architecture search at scale

TorchPQ is a python library for Approximate Nearest Neighbor Search (ANNS) and Maximum Inner Product Search (MIPS) on GPU using Product Quantization (PQ) algorithm.

Densely Connected Search Space for More Flexible Neural Architecture Search (CVPR2020)

Crab is a ﬂexible, fast recommender engine for Python that integrates classic information ﬁltering recommendation algorithms in the world of scientiﬁc Python packages (numpy, scipy, matplotlib).