Python library to extract tabular data from images and scanned PDFs

Org. Account

Last update: Dec 31, 2022

Related tags

Computer Vision ocr tabular-data table-extraction image-table-recognition pdf-table-extract extracttable

Overview

ExtractTable - API to extract tabular data from images and scanned PDFs

The motivation is to make it easy for developers to extract tabular data from images or scanned PDF files without worrying about the table area, column coordinates, rotation et al.

Prerequisite

API Key: All requests to ExtractTable are authorized by an API Key. FREE credits here. The same API Key can also be used for conversions on the browser at Web Pro.

Installation

pip install -U ExtractTable

Basic Usage

Ok, enough selling. Let the ease in coding do the talk, and the output encourages you to buy credits; put that timer on and count the LOC.

from ExtractTable import ExtractTable
et_sess = ExtractTable(api_key=YOUR_API_KEY)        # Replace your VALID API Key here
print(et_sess.check_usage())        # Checks the API Key validity as well as shows associated plan usage 
table_data = et_sess.process_file(filepath=Location_of_Image_with_Tables, output_format="df")

# To process PDF, make use of pages ("1", "1,3-4", "all") params in the read_pdf function
table_data = et_sess.process_file(filepath=Location_of_PDF_with_Tables, output_format="df", pages="all")

Detailed Library Usage

The tutorial available at takes you through

1. Installation
2. Import and check version
3. Create Session & Validate API Key
    3.1 Create Session with your API Key
    3.2 Validate the Key and check the plan usage
    3.3 Check Usage Details
4. Trigger the extraction process
    4.1 Accepted Input Types
    4.2 Process an IMAGE Input
    4.3 Process a PDF Input
    4.4 Output options
    4.5 Explore session objects
5. Explore the Output
    5.1 Output Structure
    5.2 Output Details
6. Make Corrections
    6.1 Split Merged Rows
    6.2 Split Merged Columns
    6.3 Fix Decimal Format
    6.4 Fix Date Format
7. Helpful Code Snippets
    7.1 Get text data
    7.2 Table output to Excel

Woahh, as simple as that ?!

Certainly. Do you know the current ExtractTable users use it for

Bank Statement
Medical Records
Invoice Details
Tax forms
Tender Notices

Its up to you now to explore the ways.

Explore

check the complete server response of the latest job with et_sess.ServerResponse.json()

{
    "JobStatus": <string>,                              # Status of the triggered Process  @ JOB-LEVEL
    "Pages": <integer>,                                 # Number of pages processed in this request @ PAGE-LEVEL
    "Tables": [<list of key-value objects of table>     # List of all tables found @ TABLE-LEVEL
        {
            "Page": <integer>,                              ## Page number in which this table is found
            "CharacterConfidence": <float>,                 ## Accuracy of Characters recognized from the input-page
            "LayoutConfidence": <float>,                    ## Accuracy of table layout's design decision
            "TableJson": <dict>,                            ## Table Cell Text in key-value format with index orientation - {row#: {col#: <str>}}
            "TableCoordinates": <dict>,                     ## Top-left & Bottom-right Cell Coordinates - {row#: {col#: <list(x1,y1,x2,y2)>}}
            "TableConfidence": <dict>                       ## Cell level accuracy of detected characters - {row#: {col#: <float>}}
        },
    {...}                                               ## ... more "Tables" objects
    ],
    "Lines": [<list of key-value objects>               # Pagewise Line details @ PAGE-LEVEL
        {
            "Page": <integer>,                          # Page number in which the lines are found
            "CharacterConfidence": <float>,             # Average Accuracy of all Characters recognized from the input-page
            "LinesArray": [
                <list of key-value objects of line>     # Ordered list of lines in this page @ LINE-LEVEL
                {
                    "Line": <str>,                          ## Detected text of the complete line
                    "WordsArray": [
                        <list of key-value objects>         ## Word level datails in this line @ WORD-LEVEL
                        {
                            "Conf": <float>,                    ### Accuracy of recognized characters of the word
                            "Word": <str>,                      ### Detected text of the word
                            "Loc": [x1, y1, x2, y2]             ### Top-left & Bottom-right coordinates, w.r.t the input-page width-height dimensions
                        },
                    {...}                                   ### More "WordsArray" objects
                    ]
                },
            {...}                                       ## More "LinesArray" objects
            ]
        },
    {...}                                               # More Pagewise "Lines" details
    ]
}

Bug Reports

Bug reports/fixes are most welcome and greatly appreciated with API credits. For support reach us at [email protected]

License

This project is licensed under the Apache License 2.0, see the LICENSE file for details.

Social Media

Comments

bug: holding when the program running after some samples

Describe the bug A clear and concise description of what the bug is. keep holding my apI key prefix is o6No6aqYRhrQ2MWxtDDyTeHiiUg****

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen.

Additional context Add any other context about the problem here.
bug

opened by franztao 5
bug: function "et_sess.save_output(output_folder, output_format="csv")" output file, the file name lack some alpha of the origin full name

Describe the bug A clear and concise description of what the bug is. my picture name is all suffix png. such as "[email protected]_14-1-4.png"

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen.

Additional context Add any other context about the problem here.
bug

opened by franztao 3
found some bugs and list the bugs out

Describe the bug A clear and concise description of what the bug is. 1.不能识别出垮列的文本，识别成表格时，不符合逻辑的分开成两边

2.不能识别加减号,can not recognize Plus minus sign. 31.2 + 4.98 3.不能够识别上下标，can not recognize subscript and supscript. 4.ocr识别丢失字符 loss some recognized tokens 5.长的表格，有部分没有识别出来 long size table,can not recognize the bottem part 6.cell中有化学式的，识别不出来,when there is chemical formulate in cell, can not recognize the table

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen. I can solve these problems with us.

Additional context Add any other context about the problem here.
bug

opened by franztao 2
question: what meaning is LayoutConfidence?

"CharacterConfidence": , # Average Accuracy of all Characters recognized from the input-page "LayoutConfidence": , ## Accuracy of table layout's design decision please give out the detaild decription or calculate function code about CharacterConfidence,LayoutConfidence
good first issue

opened by franztao 2
Invalid cross-device link
Describe the bug On some OS, we can not save output file to temporary directory (let's say /tmp) and move it to a new place. It throws the following error :

os.replace(each_tbl_path, os.path.join(output_folder, input_fname+os.path.basename(each_tbl_path))) OSError: [Errno 18] Invalid cross-device link: '/tmp/tmp7hqcm0fh/_table_1.csv' -> '/var/www/python/app/tmp/details_table_1.csv'

After checking the source code, it appears ExtractTable use os.replace to move the file. This method does not support moving file from a partition to an other : https://stackoverflow.com/questions/42392600/oserror-errno-18-invalid-cross-device-link

To Reproduce I use Python 3.6 in a venv. You will need two different system parts, and invoke save_output from ExtractTable-py library, to save file from a filesystem to an other. I have not tried, but I think you can simply reproduce this bug by invoking os.replace without calling ExtractTable-py.

Expected behavior Move the file from a filesystem to an other. I think using shutil.move would be a preferable way to achieve file moving than os.replace.
bug
opened by Elegye 2
MakeCorrections API - How do you chain corrections

Hi there, I'm trying to use multiple correction commands but it isn't working as the object becomes a list after the first correction. Is there something I'm missing here? Thanks!
good first issue

opened by kylebutts 1
character ocr can support latex format?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by franztao 1
please, do you have tools of transform ExtracTable output file type to CoCo file type(other open source Detection file type)?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by franztao 1
Custom output path when the output_format is csv

Is your feature request related to a problem? Please describe. When the output_format is set to csv the csv file is written to some random path in /tmp location.

Describe the solution you'd like [optional, but helpful] Define a parameter in the process_file like output_file which takes the absolute path where the file needs to be written along with the file name

opened by padmano 1
Is it possible to get the data in excel by maintaining table structure?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by jcthink 1

Character and Layout Confidence

Hi, need some definition material for Character and Layout Confidence like how it is calculated mathematically using below code. Thanks.

for idx, each_table in enumerate(et_sess.ServerResponse.json()['Tables']):
    print("CharacterConfidence = ", each_table['CharacterConfidence'])
    print("LayoutConfidence = ", each_table['LayoutConfidence'])

good first issue

opened by muhdzubair 1

Consider user hints on the table structure information

Is your feature request related to a problem? Please describe. "while you do whatever you want, why not consider the our hints" is the developers feedback on many instances

Describe alternatives you've considered Developers are tackling with their custom post processing.

Describe the solution you'd like [optional, but helpful] Pros: May be it is a worth taking a look as most of the post processing involves in similar approaches that resolves majority issues. Cons: computing cost
feature/idea

opened by akshowhini 0
Capture Vertically center aligned columns

Refer: https://stackoverflow.com/questions/58238981/extracting-table-from-a-pdf-table-without-vertical-lines

Do not miss: Joelgeraci's comment to the question
feature/idea

opened by akshowhini 0

Releases(v2.4.0)

v2.4.0(Jul 18, 2022)

Use corrections.save_output() to save the output to a folder
Source code(tar.gz)
Source code(zip)
ExtractTable-2.4.0-py3-none-any.whl(18.90 KB)
ExtractTable-2.4.0.tar.gz(16.16 KB)
v2.3.1(May 6, 2022)
Fix processing splitted PDFs

Support downloading BigFile

Source code(tar.gz)
Source code(zip)
ExtractTable-2.3.1-py3-none-any.whl(18.38 KB)
ExtractTable-2.3.1.tar.gz(16.01 KB)
v2.2.0(Apr 20, 2021)
View your transactions processed in the last 24 hours

Give user the ability to make character error corrections

Source code(tar.gz)
Source code(zip)
ExtractTable-2.2.0-py3-none-any.whl(20.52 KB)
v2.1.2(Nov 6, 2020)

To provide user control on whether to output the row & column numbers in the output file
Source code(tar.gz)
Source code(zip)
ExtractTable-2.1.2-py3-none-any.whl(20.13 KB)
ExtractTable-2.1.2.tar.gz(12.04 KB)
v2.1.0(Aug 27, 2020)
Data Cleaning on the server output made easy with MakeCorrections class.

split_merged_rows

split_merged_columns

fix_decimal_format

fix_date_format functionalities

added server_response attribute to the session for easy reference

Update Google Colab Tutorial

Save tables to multiple sheets of a single excel file

save_output functionality in session to save Tables & Text output to local

Updated Tutorial in example-code.ipynb
Source code(tar.gz)
Source code(zip)
ExtractTable-2.1.0-py3-none-any.whl(20.12 KB)
ExtractTable-2.1.0.tar.gz(12.03 KB)
v2.0.2(Jul 4, 2020)

To handle Invalid Object Exception when processing big files
Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.2-py3-none-any.whl(17.28 KB)
ExtractTable-2.0.2.tar.gz(9.27 KB)
v2.0.1(Jul 3, 2020)
#28

Maintain column and row indices order

Display JobId & Wait message for async transactions

Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.1-py3-none-any.whl(17.27 KB)
ExtractTable-2.0.1.tar.gz(9.27 KB)
v2.0.0(Apr 30, 2020)

To support the below added features at API level • Tables + Text Data: non-tabular text along with the tabular data (when tables undetected, by default, response gets text data) • Text Accuracy Details: page level character accuracy details • Cell / Word Coordinates: x,y coordinates of all words and table's cell data • Cell / Word Level Accuracy: word level accuracy details • Non-English characters: for non-english alphabets like Mandrin, Japanese etc
Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.0-py3-none-any.whl(17.04 KB)
ExtractTable-2.0.0.tar.gz(9.09 KB)
v1.2.1.2(Dec 1, 2019)

Fixed Columns were not in order in the output
Source code(tar.gz)
Source code(zip)
ExtractTable-1.2.1.2-py3-none-any.whl(14.09 KB)
ExtractTable-1.2.1.2.tar.gz(8.01 KB)
v1.1.0(Oct 20, 2019)

Support URL as input which downloads the file to the temporary directory for the instance, auto deleted on success
Source code(tar.gz)
Source code(zip)
v1.0.1(Oct 7, 2019)

The first stable library to make it easier for python developers to use ExtractTable's API to extract tabular data (table) from images and scanned PDFs without worrying about table area, column regions, image rotation et al.
Source code(tar.gz)
Source code(zip)
ExtractTable-1.0.1-py3-none-any.whl(12.42 KB)
ExtractTable-1.0.1.tar.gz(6.54 KB)

Owner

Org. Account

You, I and they have the same problem to solve !?!?

GitHub https://extracttable.com

Unofficial implementation of "TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images"

TableNet Unofficial implementation of ICDAR 2019 paper : TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from

243 Dec 30, 2022

Detect text blocks and OCR poorly scanned PDFs in bulk. Python module available via pip.

doc2text doc2text extracts higher quality text by fixing common scan errors Developing text corpora can be a massive pain in the butt. Much of the tex

1.3k Jan 4, 2023

This is a GUI for scrapping PDFs with the help of optical character recognition making easier than ever to scrape PDFs.

pdf-scraper-with-ocr With this tool I am aiming to facilitate the work of those who need to scrape PDFs either by hand or using tools that doesn't imp

75 Oct 21, 2022

Library used to deskew a scanned document

Deskew //Note: Skew is measured in degrees. Deskewing is a process whereby skew is removed by rotating an image by the same amount as its skew but in

273 Jan 6, 2023

A post-processing tool for scanned sheets of paper.

unpaper Originally written by Jens Gulden — see AUTHORS for more information. Licensed under GNU GPL v2 — see COPYING for more information. Overview u

27 Dec 7, 2022

Deskew is a command line tool for deskewing scanned text documents. It uses Hough transform to detect "text lines" in the image. As an output, you get an image rotated so that the lines are horizontal.

Deskew by Marek Mauder https://galfar.vevb.net/deskew https://github.com/galfar/deskew v1.30 2019-06-07 Overview Deskew is a command line tool for des

127 Dec 3, 2022

A tool for extracting text from scanned documents (via OCR), with user-defined post-processing.

The project is based on older versions of tesseract and other tools, and is now superseded by another project which allows for more granular control o

32 Jul 24, 2022

Some bits of javascript to transcribe scanned pages using PageXML

nashi (nasḫī) Some bits of javascript to transcribe scanned pages using PageXML. Both ltr and rtl languages are supported. Try it! But wait, there's m

15 Nov 9, 2022

scantailor - Scan Tailor is an interactive post-processing tool for scanned pages.

Scan Tailor - scantailor.org This project is no longer maintained, and has not been maintained for a while. About Scan Tailor is an interactive post-p

1.5k Dec 28, 2022

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted. ocrmypdf # it's a scriptable c

7.9k Jan 3, 2023

Recognizing the text contents from a scanned visiting card

Recognizing the text contents from a scanned visiting card. The application which is used to recognize the text from scanned images,printeddocuments,r

1 Jan 28, 2022

A bot that extract text from images using the Tesseract OCR.

Text from image (OCR) @ocr_text_bot A simple bot to extract text from images. Usage What do I need? A AWS key configured locally, see here. NodeJS. I

4 Aug 6, 2021

Turn images of tables into CSV data. Detect tables from images and run OCR on the cells.

Table of Contents Overview Requirements Demo Modules Overview This python package contains modules to help with finding and extracting tabular data fr

311 Dec 24, 2022

Code for generating synthetic text images as described in "Synthetic Data for Text Localisation in Natural Images", Ankush Gupta, Andrea Vedaldi, Andrew Zisserman, CVPR 2016.

SynthText Code for generating synthetic text images as described in "Synthetic Data for Text Localisation in Natural Images", Ankush Gupta, Andrea Ved

1.8k Dec 28, 2022

Scan the MRZ code of a passport and extract the firstname, lastname, passport number, nationality, date of birth, expiration date and personal numer.

PassportScanner Works with 2 and 3 line identity documents. What is this With PassportScanner you can use your camera to scan the MRZ code of a passpo

441 Dec 24, 2022

Let's explore how we can extract text from forms

Form Segmentation Let's explore how we can extract text from any forms / scanned pages. Objectives The goal is to find an algorithm that can extract t

42 Jun 5, 2022

IMGUR5K handwriting set. It is a handwritten in-the-wild dataset, which contains challenging real world handwritten samples from different writers.The dataset is shared as a set of image urls with annotations. This code downloads the images and verifies the hash to the image to avoid data contamination.

IMGUR5K Handwriting Dataset To run the code for downloading the urls and generate corresponding annotations : Usage: python download_imgur5k.py --data

213 Dec 26, 2022

Basic functions manipulating images using the OpenCV library

OpenCV Basic functions manipulating images using the OpenCV library. Reading Ima

3 Feb 17, 2022

Genalog is an open source, cross-platform python package allowing generation of synthetic document images with custom degradations and text alignment capabilities.

235 Dec 22, 2022

Python library to extract tabular data from images and scanned PDFs

Related tags

Overview

Overview

Prerequisite

Installation

Basic Usage

Detailed Library Usage

Woahh, as simple as that ?!

Explore

Bug Reports

License

Social Media

Comments

Releases(v2.4.0)

v2.4.0(Jul 18, 2022)

v2.3.1(May 6, 2022)

v2.2.0(Apr 20, 2021)

v2.1.2(Nov 6, 2020)

v2.1.0(Aug 27, 2020)

v2.0.2(Jul 4, 2020)

v2.0.1(Jul 3, 2020)

v2.0.0(Apr 30, 2020)

v1.2.1.2(Dec 1, 2019)

v1.1.0(Oct 20, 2019)

v1.0.1(Oct 7, 2019)

Owner

Org. Account

Unofficial implementation of "TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images"

Detect text blocks and OCR poorly scanned PDFs in bulk. Python module available via pip.

This is a GUI for scrapping PDFs with the help of optical character recognition making easier than ever to scrape PDFs.

Library used to deskew a scanned document

A post-processing tool for scanned sheets of paper.

Deskew is a command line tool for deskewing scanned text documents. It uses Hough transform to detect "text lines" in the image. As an output, you get an image rotated so that the lines are horizontal.

A tool for extracting text from scanned documents (via OCR), with user-defined post-processing.

Some bits of javascript to transcribe scanned pages using PageXML

scantailor - Scan Tailor is an interactive post-processing tool for scanned pages.

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

Recognizing the text contents from a scanned visiting card

A bot that extract text from images using the Tesseract OCR.

Turn images of tables into CSV data. Detect tables from images and run OCR on the cells.

Code for generating synthetic text images as described in "Synthetic Data for Text Localisation in Natural Images", Ankush Gupta, Andrea Vedaldi, Andrew Zisserman, CVPR 2016.

Scan the MRZ code of a passport and extract the firstname, lastname, passport number, nationality, date of birth, expiration date and personal numer.

Let's explore how we can extract text from forms

Basic functions manipulating images using the OpenCV library

Genalog is an open source, cross-platform python package allowing generation of synthetic document images with custom degradations and text alignment capabilities.