Bump pypdf2 from 1.25.1 to 2.1.0 by dependabot[bot] · Pull Request #392 · SpamExperts/OrangeAssassin

dependabot · 2022-06-07T05:10:35Z

Bumps pypdf2 from 1.25.1 to 2.1.0.

Release notes

Version 2.1.0, 2022-06-06

What's Changed

The highlight of the 2.1.0 release is the most massive improvement to the text extraction capabilities of PyPDF2 since 2016 🥳🎊 A very big thank you goes to pubpub-zz who took a lot of time and knowledge about the PDF format to finally get those improvements into PyPDF2. Thank you 🤗💚

In case the new function causes any issues, you can use _extract_text_old for the old functionality. Please also open a bug ticket in that case.

There were several people who have attempted to bring similar improvements to PyPDF2. All of those were valuable. The main reason why they didn't get merged is the big amount of open PRs / issues. pubpub-zz was the most comprehensive PR which also incorporated the latest changes of PyPDF2 2.0.0.

Thank you to VictorCarlquist for #858 and asabramo for #464 🤗

New Features (ENH)

Massive text extraction improvement (#924). Closed many open issues:

Exceptions / missing spaces in extract_text() method (#17) 🕺

Whitespace issues in extract_text() (#42) 💃

pypdf2 reads the hifenated words in a new line (#246)

PyPDF2 failing to read unicode character (#37)

Unable to read bullets (#230)

ExtractText yields nothing for apparently good PDF (#168) 🎉

Encoding issue in extract_text() (#235)

extractText() doesn't work on Chinese PDF (#252)

encoding error (#260)

Trouble with apostophes in names in text "O'Doul" (#384)

extract_text works for some PDF files, but not the others (#437)

Euro sign not being recognized by extractText (#443)

Failed extracting text from French texts (#524)

extract_text doesn't extract ligatures correctly (#598)

reading spanish text - mark convert issue (#635)

Read PDF changed from text to random symbols (#654)

.extractText() reads / as 1. (#789)

Update glyphlist (#947) - inspired by #464

Allow adding PageRange objects (#948)

Bug Fixes (BUG)

Delete .python-version file (#944)

Compare StreamObject.decoded_self with None (#931)

Robustness (ROB)

Fix some conversion errors on non conform PDF (#932)

Documentation (DOC)

Elaborate on PDF text extraction difficulties (#939)

... (truncated)

Changelog

Sourced from pypdf2's changelog.

Version 2.1.0, 2022-06-06

The highlight of the 2.1.0 release is the most massive improvement to the text extraction capabilities of PyPDF2 since 2016 🥳🎊 A very big thank you goes to pubpub-zz who took a lot of time and knowledge about the PDF format to finally get those improvements into PyPDF2. Thank you 🤗💚

In case the new function causes any issues, you can use _extract_text_old for the old functionality. Please also open a bug ticket in that case.

There were several people who have attempted to bring similar improvements to PyPDF2. All of those were valuable. The main reason why they didn't get merged is the big amount of open PRs / issues. pubpub-zz was the most comprehensive PR which also incorporated the latest changes of PyPDF2 2.0.0.

Thank you to VictorCarlquist for #858 and asabramo for #464 🤗

New Features (ENH):

Massive text extraction improvement (#924). Closed many open issues:

Exceptions / missing spaces in extract_text() method (#17) 🕺

Whitespace issues in extract_text() (#42) 💃

pypdf2 reads the hifenated words in a new line (#246)

PyPDF2 failing to read unicode character (#37)

Unable to read bullets (#230)

ExtractText yields nothing for apparently good PDF (#168) 🎉

Encoding issue in extract_text() (#235)

extractText() doesn't work on Chinese PDF (#252)

encoding error (#260)

Trouble with apostophes in names in text "O'Doul" (#384)

extract_text works for some PDF files, but not the others (#437)

Euro sign not being recognized by extractText (#443)

Failed extracting text from French texts (#524)

extract_text doesn't extract ligatures correctly (#598)

reading spanish text - mark convert issue (#635)

Read PDF changed from text to random symbols (#654)

.extractText() reads / as 1. (#789)

Update glyphlist (#947) - inspired by #464

Allow adding PageRange objects (#948)

Bug Fixes (BUG):

Delete .python-version file (#944)

Compare StreamObject.decoded_self with None (#931)

Robustness (ROB):

Fix some conversion errors on non conform PDF (#932)

Documentation (DOC):

... (truncated)

Commits

4e44122 REL: 2.1.0
babe32e TST: Text extraction for non-latin alphabets (#954)
2a1db78 ENH: Allow adding PageRange objects (#948)
4baedb2 STY: black, isort, Flake8, splitting buildCharMap (#950)
b008412 TST: Ignore PdfReadWarning in benchmark (#949)
3a0bd5e ENH: Update glyphlist (#947)
0b087dc DOC: Elaborate on PDF text extraction difficulties (#939)
81a9987 ROB: Fix some conversion errors on non conform PDF (#932)
648e308 ENH: Massive text extraction improvement (#924)
1df859c TST: writer.remove_text (#946)
Additional commits viewable in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.

Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

@dependabot rebase will rebase this PR
@dependabot recreate will recreate this PR, overwriting any edits that have been made to it
@dependabot merge will merge this PR after your CI passes on it
@dependabot squash and merge will squash and merge this PR after your CI passes on it
@dependabot cancel merge will cancel a previously requested merge and block automerging
@dependabot reopen will reopen this PR if it is closed
@dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually
@dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
@dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
@dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Bumps [pypdf2](https://github.com/py-pdf/PyPDF2) from 1.25.1 to 2.1.0. - [Release notes](https://github.com/py-pdf/PyPDF2/releases) - [Changelog](https://github.com/py-pdf/PyPDF2/blob/main/CHANGELOG) - [Commits](py-pdf/pypdf@v1.25.1...2.1.0) --- updated-dependencies: - dependency-name: pypdf2 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>

dependabot · 2022-06-13T05:13:14Z

Superseded by #393.

dependabot bot added the dependencies Pull requests that update a dependency file label Jun 7, 2022

dependabot bot mentioned this pull request Jun 7, 2022

Bump pypdf2 from 1.25.1 to 2.0.0 #389

Closed

dependabot bot closed this Jun 13, 2022

dependabot bot deleted the dependabot/pip/pypdf2-2.1.0 branch June 13, 2022 05:13

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Comments

Bump pypdf2 from 1.25.1 to 2.1.0#392

Bump pypdf2 from 1.25.1 to 2.1.0#392
dependabot[bot] wants to merge 1 commit intomasterfrom
dependabot/pip/pypdf2-2.1.0

dependabot bot commented on behalf of github Jun 7, 2022

Uh oh!

dependabot bot commented on behalf of github Jun 13, 2022

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

0 participants

Comments

Conversation

dependabot bot commented on behalf of github Jun 7, 2022

Version 2.1.0, 2022-06-06

What's Changed

New Features (ENH)

Bug Fixes (BUG)

Robustness (ROB)

Documentation (DOC)

Version 2.1.0, 2022-06-06

Uh oh!

dependabot bot commented on behalf of github Jun 13, 2022

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

0 participants