PicAxe: Extracting Figures from Structurally and Syntactically Heterogeneous Corpora of PDF Files
PicAxe is open-source Python software that researchers can use to extract figures from corpora of PDF files that contain text and images. It is designed to extract figures from corpora that include both scanned and “born-digital” PDF files (structurally heterogeneous) of documents from different cul...
Shranjeno v:
| Principais autores: | , , , , , |
|---|---|
| Format: | Artigo |
| Jezik: | Inglês |
| Izdano: |
Ubiquity Press
2025-12-01
|
| Serija: | Journal of Open Research Software |
| Teme: | |
| Online dostop: | https://account.openresearchsoftware.metajnl.com/index.php/up-j-jors/article/view/574 |
| Oznake: |
Brez oznak, prvi označite!
|
