PDFParser

A pure Swift library for extracting text information from pdf files, such as text blocks with coordinates and font information. Also includes a true type font parser for glyph width computation.

Parsing code based on PDFKitten https://github.com/KurtCode/PDFKitten TrueType parser based on http://stevehanov.ca/blog/index.php?id=143

Parsing is done very simply, and returns TextBlocks structs, that can be later indexed by custom code. A simple indexer is provided, assuming single column layout, aggregating words.

var documentIndexer = SimpleDocumentIndexer()
let documentPath = Bundle.main.path(forResource: "Kurt the Cat", ofType: "pdf", inDirectory: nil, forLocalization: nil)

let parser = try! Parser(documentURL: URL(fileURLWithPath: documentPath!), delegate:self, indexer: documentIndexer)
parser.parse()

print( "All Text Blocks Raw dump : \n")
print(documentIndexer.pageIndexes[1]!.textBlocks)

print( "\nWords per lines : \n")
print(documentIndexer.pageIndexes[1]!.allLinesDescription())

ViewController in the DemoApp displays UILabel for textblocks. This lets you see if the frames for the textblock returned by the parser is correct.

This code is not ready for production. Use at your own risk. This code is probably way too unoptimized to be used for anything latency-sensitive. It was meant to be easy to understand and correct first and foremost.

Name		Name	Last commit message	Last commit date
Latest commit History 14 Commits
PDFParser.xcworkspace		PDFParser.xcworkspace
PDFParser		PDFParser
.gitignore		.gitignore
README.md		README.md

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

PDFParser.xcworkspace

PDFParser.xcworkspace

PDFParser

PDFParser

.gitignore

.gitignore

README.md

README.md

Repository files navigation

PDFParser

About

Releases

Packages

Languages

SimpleApp/PDFParser

Folders and files

Latest commit

History

Repository files navigation

PDFParser

About

Topics

Resources

Stars

Watchers

Forks

Languages