Metadata-Version: 2.4 Name: unstructured Version: 0.18.27 Summary: A library that prepares raw documents for downstream ML tasks. Home-page: https://github.com/Unstructured-IO/unstructured Author: Unstructured Technologies Author-email: devops@unstructuredai.io License: Apache-2.0 Keywords: NLP PDF HTML CV XML parsing preprocessing Classifier: Development Status :: 4 - Beta Classifier: Intended Audience :: Developers Classifier: Intended Audience :: Education Classifier: Intended Audience :: Science/Research Classifier: License :: OSI Approved :: Apache Software License Classifier: Operating System :: OS Independent Classifier: Programming Language :: Python :: 3 Classifier: Programming Language :: Python :: 3.10 Classifier: Programming Language :: Python :: 3.11 Classifier: Programming Language :: Python :: 3.12 Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence Requires-Python: >=3.10.0 Description-Content-Type: text/markdown License-File: LICENSE.md Requires-Dist: charset-normalizer Requires-Dist: filetype Requires-Dist: python-magic Requires-Dist: lxml Requires-Dist: nltk Requires-Dist: requests Requires-Dist: beautifulsoup4 Requires-Dist: emoji Requires-Dist: dataclasses-json Requires-Dist: python-iso639 Requires-Dist: langdetect Requires-Dist: numpy Requires-Dist: rapidfuzz Requires-Dist: backoff Requires-Dist: typing-extensions Requires-Dist: unstructured-client Requires-Dist: wrapt Requires-Dist: tqdm Requires-Dist: psutil Requires-Dist: python-oxmsg Requires-Dist: html5lib Provides-Extra: all-docs Requires-Dist: markdown; extra == "all-docs" Requires-Dist: openpyxl; extra == "all-docs" Requires-Dist: python-pptx>=1.0.1; extra == "all-docs" Requires-Dist: python-docx>=1.1.2; extra == "all-docs" Requires-Dist: pypdf; extra == "all-docs" Requires-Dist: onnx>=1.17.0; extra == "all-docs" Requires-Dist: google-cloud-vision; extra == "all-docs" Requires-Dist: pi_heif; extra == "all-docs" Requires-Dist: pdf2image; extra == "all-docs" Requires-Dist: effdet; extra == "all-docs" Requires-Dist: networkx; extra == "all-docs" Requires-Dist: pdfminer.six; extra == "all-docs" Requires-Dist: xlrd; extra == "all-docs" Requires-Dist: onnxruntime>=1.19.0; extra == "all-docs" Requires-Dist: unstructured.pytesseract>=0.3.12; extra == "all-docs" Requires-Dist: pandas; extra == "all-docs" Requires-Dist: pikepdf; extra == "all-docs" Requires-Dist: unstructured-inference>=1.1.1; extra == "all-docs" Requires-Dist: msoffcrypto-tool; extra == "all-docs" Requires-Dist: pypandoc; extra == "all-docs" Provides-Extra: csv Requires-Dist: pandas; extra == "csv" Provides-Extra: doc Requires-Dist: python-docx>=1.1.2; extra == "doc" Provides-Extra: docx Requires-Dist: python-docx>=1.1.2; extra == "docx" Provides-Extra: epub Requires-Dist: pypandoc; extra == "epub" Provides-Extra: image Requires-Dist: onnx>=1.17.0; extra == "image" Requires-Dist: onnxruntime>=1.19.0; extra == "image" Requires-Dist: pdf2image; extra == "image" Requires-Dist: pdfminer.six; extra == "image" Requires-Dist: pikepdf; extra == "image" Requires-Dist: pi_heif; extra == "image" Requires-Dist: pypdf; extra == "image" Requires-Dist: google-cloud-vision; extra == "image" Requires-Dist: effdet; extra == "image" Requires-Dist: unstructured-inference>=1.1.1; extra == "image" Requires-Dist: unstructured.pytesseract>=0.3.12; extra == "image" Provides-Extra: md Requires-Dist: markdown; extra == "md" Provides-Extra: odt Requires-Dist: python-docx>=1.1.2; extra == "odt" Requires-Dist: pypandoc; extra == "odt" Provides-Extra: org Requires-Dist: pypandoc; extra == "org" Provides-Extra: pdf Requires-Dist: onnx>=1.17.0; extra == "pdf" Requires-Dist: onnxruntime>=1.19.0; extra == "pdf" Requires-Dist: pdf2image; extra == "pdf" Requires-Dist: pdfminer.six; extra == "pdf" Requires-Dist: pikepdf; extra == "pdf" Requires-Dist: pi_heif; extra == "pdf" Requires-Dist: pypdf; extra == "pdf" Requires-Dist: google-cloud-vision; extra == "pdf" Requires-Dist: effdet; extra == "pdf" Requires-Dist: unstructured-inference>=1.1.1; extra == "pdf" Requires-Dist: unstructured.pytesseract>=0.3.12; extra == "pdf" Provides-Extra: ppt Requires-Dist: python-pptx>=1.0.1; extra == "ppt" Provides-Extra: pptx Requires-Dist: python-pptx>=1.0.1; extra == "pptx" Provides-Extra: rtf Requires-Dist: pypandoc; extra == "rtf" Provides-Extra: rst Requires-Dist: pypandoc; extra == "rst" Provides-Extra: tsv Requires-Dist: pandas; extra == "tsv" Provides-Extra: xlsx Requires-Dist: openpyxl; extra == "xlsx" Requires-Dist: pandas; extra == "xlsx" Requires-Dist: xlrd; extra == "xlsx" Requires-Dist: networkx; extra == "xlsx" Requires-Dist: msoffcrypto-tool; extra == "xlsx" Provides-Extra: huggingface Requires-Dist: langdetect; extra == "huggingface" Requires-Dist: sacremoses; extra == "huggingface" Requires-Dist: sentencepiece; extra == "huggingface" Requires-Dist: torch; extra == "huggingface" Requires-Dist: transformers; extra == "huggingface" Provides-Extra: local-inference Requires-Dist: markdown; extra == "local-inference" Requires-Dist: openpyxl; extra == "local-inference" Requires-Dist: python-pptx>=1.0.1; extra == "local-inference" Requires-Dist: python-docx>=1.1.2; extra == "local-inference" Requires-Dist: pypdf; extra == "local-inference" Requires-Dist: onnx>=1.17.0; extra == "local-inference" Requires-Dist: google-cloud-vision; extra == "local-inference" Requires-Dist: pi_heif; extra == "local-inference" Requires-Dist: pdf2image; extra == "local-inference" Requires-Dist: effdet; extra == "local-inference" Requires-Dist: networkx; extra == "local-inference" Requires-Dist: pdfminer.six; extra == "local-inference" Requires-Dist: xlrd; extra == "local-inference" Requires-Dist: onnxruntime>=1.19.0; extra == "local-inference" Requires-Dist: unstructured.pytesseract>=0.3.12; extra == "local-inference" Requires-Dist: pandas; extra == "local-inference" Requires-Dist: pikepdf; extra == "local-inference" Requires-Dist: unstructured-inference>=1.1.1; extra == "local-inference" Requires-Dist: msoffcrypto-tool; extra == "local-inference" Requires-Dist: pypandoc; extra == "local-inference" Provides-Extra: paddleocr Requires-Dist: paddlepaddle>=3.0.0b1; extra == "paddleocr" Requires-Dist: unstructured.paddleocr==2.10.0; extra == "paddleocr" Dynamic: author Dynamic: author-email Dynamic: classifier Dynamic: description Dynamic: description-content-type Dynamic: home-page Dynamic: keywords Dynamic: license Dynamic: license-file Dynamic: provides-extra Dynamic: requires-dist Dynamic: requires-python Dynamic: summary
Open-Source Pre-Processing Tools for Unstructured Data
The `unstructured` library provides open-source components for ingesting and pre-processing images and text documents, such as PDFs, HTML, Word docs, and [many more](https://docs.unstructured.io/open-source/core-functionality/partitioning). The use cases of `unstructured` revolve around streamlining and optimizing the data processing workflow for LLMs. `unstructured` modular functions and connectors form a cohesive system that simplifies data ingestion and pre-processing, making it adaptable to different platforms and efficient in transforming unstructured data into structured outputs. ## Try the Unstructured Platform Product Ready to move your data processing pipeline to production, and take advantage of advanced features? Check out [Unstructured Platform](https://unstructured.io/enterprise). In addition to better processing performance, take advantage of chunking, embedding, and image and table enrichment generation, all from a low code UI or an API. [Request a demo](https://unstructured.io/contact) from our sales team to learn more about how to get started. ## :eight_pointed_black_star: Quick Start There are several ways to use the `unstructured` library: * [Run the library in a container](https://github.com/Unstructured-IO/unstructured#run-the-library-in-a-container) or * Install the library 1. [Install from PyPI](https://github.com/Unstructured-IO/unstructured#installing-the-library) 2. [Install for local development](https://github.com/Unstructured-IO/unstructured#installation-instructions-for-local-development) * For installation with `conda` on Windows system, please refer to the [documentation](https://unstructured-io.github.io/unstructured/installing.html#installation-with-conda-on-windows) ### Run the library in a container The following instructions are intended to help you get up and running using Docker to interact with `unstructured`. See [here](https://docs.docker.com/get-docker/) if you don't already have docker installed on your machine. NOTE: we build multi-platform images to support both x86_64 and Apple silicon hardware. `docker pull` should download the corresponding image for your architecture, but you can specify with `--platform` (e.g. `--platform linux/amd64`) if needed. We build Docker images for all pushes to `main`. We tag each image with the corresponding short commit hash (e.g. `fbc7a69`) and the application version (e.g. `0.5.5-dev1`). We also tag the most recent image with `latest`. To leverage this, `docker pull` from our image repository. ```bash docker pull downloads.unstructured.io/unstructured-io/unstructured:latest ``` Once pulled, you can create a container from this image and shell to it. ```bash # create the container docker run -dt --name unstructured downloads.unstructured.io/unstructured-io/unstructured:latest # this will drop you into a bash shell where the Docker image is running docker exec -it unstructured bash ``` You can also build your own Docker image. Note that the base image is `wolfi-base`, which is updated regularly. If you are building the image locally, it is possible `docker-build` could fail due to upstream changes in `wolfi-base`. If you only plan on parsing one type of data you can speed up building the image by commenting out some of the packages/requirements necessary for other data types. See Dockerfile to know which lines are necessary for your use case. ```bash make docker-build # this will drop you into a bash shell where the Docker image is running make docker-start-bash ``` Once in the running container, you can try things directly in Python interpreter's interactive mode. ```bash # this will drop you into a python console so you can run the below partition functions python3 >>> from unstructured.partition.pdf import partition_pdf >>> elements = partition_pdf(filename="example-docs/layout-parser-paper-fast.pdf") >>> from unstructured.partition.text import partition_text >>> elements = partition_text(filename="example-docs/fake-text.txt") ``` ### Installing the library Use the following instructions to get up and running with `unstructured` and test your installation. - Install the Python SDK to support all document types with `pip install "unstructured[all-docs]"` - For plain text files, HTML, XML, JSON and Emails that do not require any extra dependencies, you can run `pip install unstructured` - To process other doc types, you can install the extras required for those documents, such as `pip install "unstructured[docx,pptx]"` - Install the following system dependencies if they are not already available on your system. Depending on what document types you're parsing, you may not need all of these. - `libmagic-dev` (filetype detection) - `poppler-utils` (images and PDFs) - `tesseract-ocr` (images and PDFs, install `tesseract-lang` for additional language support) - `libreoffice` (MS Office docs) - `pandoc` (EPUBs, RTFs and Open Office docs). Please note that to handle RTF files, you need version `2.14.2` or newer. Running either `make install-pandoc` or `./scripts/install-pandoc.sh` will install the correct version for you. - For suggestions on how to install on the Windows and to learn about dependencies for other features, see the installation documentation [here](https://unstructured-io.github.io/unstructured/installing.html). At this point, you should be able to run the following code: ```python from unstructured.partition.auto import partition elements = partition(filename="example-docs/eml/fake-email.eml") print("\n\n".join([str(el) for el in elements])) ``` ### Installation Instructions for Local Development The following instructions are intended to help you get up and running with `unstructured` locally if you are planning to contribute to the project. * Using `pyenv` to manage virtualenv's is recommended but not necessary * Mac install instructions. See [here](https://github.com/Unstructured-IO/community#mac--homebrew) for more detailed instructions. * `brew install pyenv-virtualenv` * `pyenv install 3.10` * Linux instructions are available [here](https://github.com/Unstructured-IO/community#linux). * Create a virtualenv to work in and activate it, e.g. for one named `unstructured`: `pyenv virtualenv 3.10 unstructured`