Measure file similarity in a many-to-many fashion

GatorEducator

Last update: Feb 2, 2022

Related tags

Overview

Mesi

Mesi is a tool to measure the similarity in a many-to-many fashion of long-form documents like Python source code or technical writing. The output can be useful in determining which of a collection of files are the most similar to each other.

Installation

Python 3.9+ and pipx are recommended, although Python 3.6+ and/or pip will also work.

pipx install mesi

If you'd like to test out Mesi before installing it, use the remote execution feature of pipx, which will temporarily download Mesi and run it in an isolated virtual environment.

pipx run mesi --help

Usage

For a directory structure that looks like:

lab-one
├── StudentOne
│   ├── pyproject.toml
│   ├── deliverables
│   │   └── python_program.py
│   └── README.md
├── StudentTwo
│   ├── pyproject.toml
│   ├── deliverables
│   │   └── python_program.py
│   └── README.md
│

where similarity should be measured between each student's deliverables/python_program.py file, run the command:

mesi lab-one/*/deliverables/python_program.py

A lower distance in the produced table equates to a higher degree of similarity.

See the help menu (mesi --help) for additional options and configuration.

Algorithms

There are many algorithms to choose from when comparing string similarity! Mesi implements all the algorithms provided by TextDistance. In general levenshtein is never a bad choice, which is why it is the default.

Bugs/Requests

Please use the GitHub issue tracker to submit bugs or request new features, options, or algorithms.

Dependencies

Mesi uses two primary dependencies for text similarity calculation: polyleven, and TextDistance. Polyleven is the default, as its singular implementation of Levenshtein distance can be faster in most situations. However, if a different edit distance algorithm is requested, TextDistance's implementations will be used.

License

Distributed under the terms of the GPL v3 license, mesi is free and open source software.

File-manager - A basic file manager, written in Python

File Manager A basic file manager, written in Python. Installation Install Pytho

1 Feb 5, 2022

Two scripts help you to convert csv file to md file by template

Two scripts help you to convert csv file to md file by template. One help you generate multiple md files with different filenames from the first colume of csv file. Another can generate one md file with several blocks.

2 Oct 15, 2022

A simple Python code that takes input from a csv file and makes it into a vcf file.

Contacts-Maker A simple Python code that takes input from a csv file and makes it into a vcf file. Imagine a college or a large community where each y

1 Feb 13, 2022

This program can help you to move and rename many files at once

This program can help you to rename and save many files in a folder in seconds, but don't give the same name to files, it can delete both files.

1 Oct 10, 2022

Object-oriented file system path manipulation

path (aka path pie, formerly path.py) implements path objects as first-class entities, allowing common operations on files to be invoked on those path

1k Dec 28, 2022

An object-oriented approach to Python file/directory operations.

Unipath An object-oriented approach to file/directory operations Version: 1.1 Home page: https://github.com/mikeorr/Unipath Docs: https://github.com/m

506 Dec 29, 2022

File support for asyncio

aiofiles: file support for asyncio aiofiles is an Apache2 licensed library, written in Python, for handling local disk files in asyncio applications.

2.1k Jan 1, 2023

Object-oriented file system path manipulation

path (aka path pie, formerly path.py) implements path objects as first-class entities, allowing common operations on files to be invoked on those path

1k Dec 28, 2022

A platform independent file lock for Python

py-filelock This package contains a single module, which implements a platform independent file lock in Python, which provides a simple way of inter-p

497 Jan 5, 2023

Releases(v1.1.0)

v1.1.0(Dec 8, 2021)

This release adds an --average option, which prints the average distance computed against the given files after the individual comparison distances. An unimplemented --distribution option, intended for use in the future, was also added.
Source code(tar.gz)
Source code(zip)
v1.0.2(Oct 26, 2021)

This release adds more documentation, a --version option, and removes the bright white table styling from previous versions that were hard to read on light-themed terminals.

Full Changelog: https://github.com/Michionlion/mesi/compare/v1.0.1...v1.0.2
Source code(tar.gz)
Source code(zip)
v1.0.1(Oct 26, 2021)

This release fixes a bug where in some situations, some distinct parts of compared file names were not displayed. Additionally, a new option, --table-format, was added. This allows configuration of the output table format, using tabulate's table formats.

Full Changelog: https://github.com/Michionlion/mesi/compare/v1.0.0...v1.0.1
Source code(tar.gz)
Source code(zip)
v1.0.0(Oct 25, 2021)
Mesi is a tool to check the similarity between many different files; this is its first release, with support for many basic features and lots of text similarity/distance algorithms. You can even try out Mesi without downloading it!

pipx run mesi --help

If you don't have pipx installed, you may need to install it, but that's something you should do anyways. :smile:
Source code(tar.gz)
Source code(zip)

Owner

GatorEducator

Software tools developed at Allegheny College for computer science courses

GitHub

gitfs is a FUSE file system that fully integrates with git - Version controlled file system

gitfs is a FUSE file system that fully integrates with git. You can mount a remote repository's branch locally, and any subsequent changes made to the files will be automatically committed to the remote.

2.3k Jan 8, 2023

A python script to convert an ucompressed Gnucash XML file to a text file for Ledger and hledger.

README 1 gnucash2ledger gnucash2ledger is a Python script based on the Github Gist by nonducor (nonducor/gcash2ledger.py). This Python script will tak

0 Jan 28, 2022

This is a file deletion program that asks you for an extension of a file (.mp3, .pdf, .docx, etc.) to delete all of the files in a dir that have that extension.

FileBulk This is a file deletion program that asks you for an extension of a file (.mp3, .pdf, .docx, etc.) to delete all of the files in a dir that h

1 Jun 26, 2022

Measure file similarity in a many-to-many fashion

Related tags

Overview

Mesi

Installation

Usage

Algorithms

Bugs/Requests

Dependencies

License

You might also like...

File-manager - A basic file manager, written in Python

Two scripts help you to convert csv file to md file by template

A simple Python code that takes input from a csv file and makes it into a vcf file.

This program can help you to move and rename many files at once

Object-oriented file system path manipulation

An object-oriented approach to Python file/directory operations.

File support for asyncio

Object-oriented file system path manipulation

A platform independent file lock for Python

Releases(v1.1.0)

v1.1.0(Dec 8, 2021)

v1.0.2(Oct 26, 2021)

v1.0.1(Oct 26, 2021)

v1.0.0(Oct 25, 2021)

Owner

GatorEducator

gitfs is a FUSE file system that fully integrates with git - Version controlled file system

A python script to convert an ucompressed Gnucash XML file to a text file for Ledger and hledger.

This is a file deletion program that asks you for an extension of a file (.mp3, .pdf, .docx, etc.) to delete all of the files in a dir that have that extension.

Python package to read and display segregated file names present in a directory based on type of the file

Search for files under the specified directory. Extract the file name and file path and import them as data.

Small-File-Explorer - I coded a small file explorer with several options

Pti-file-format - Reverse engineering the Polyend Tracker instrument file format

Generates a clean .txt file of contents of a 3 lined csv file

PaddingZip - a tool that you can craft a zip file that contains the padding characters between the file content.

Extract longest transcript or longest CDS transcript from GTF annotation file or gencode transcripts fasta file.