a
Re @ s0 d dl Z d dlZddlmZ G dd deZdS ) N )ProbingStatec @ sn e Zd ZdZdddZdd Zedd Zd d
Zedd Z d
d Z
edd Zedd Z
edd ZdS )
CharSetProbergffffff?Nc C s d | _ || _tt| _d S N)_statelang_filterlogging getLogger__name__logger)selfr r
/builddir/build/BUILDROOT/alt-python39-pip-21.3.1-2.el8.x86_64/opt/alt/python39/lib/python3.9/site-packages/pip/_vendor/chardet/charsetprober.py__init__' s zCharSetProber.__init__c C s t j| _d S r )r DETECTINGr r r
r
r reset, s zCharSetProber.resetc C s d S r r
r r
r
r charset_name/ s zCharSetProber.charset_namec C s d S r r
)r bufr
r
r feed3 s zCharSetProber.feedc C s | j S r )r r r
r
r state6 s zCharSetProber.statec C s dS )Ng r
r r
r
r get_confidence: s zCharSetProber.get_confidencec C s t dd| } | S )Ns ([ -])+ )resub)r r
r
r filter_high_byte_only= s z#CharSetProber.filter_high_byte_onlyc C s\ t }td| }|D ]@}||dd |dd }| sL|dk rLd}|| q|S )u9
We define three types of bytes:
alphabet: english alphabets [a-zA-Z]
international: international characters [-ÿ]
marker: everything else [^a-zA-Z-ÿ]
The input buffer can be thought to contain a series of words delimited
by markers. This function works to filter all words that contain at
least one international character. All contiguous sequences of markers
are replaced by a single space ascii character.
This filter applies to all scripts which do not use English characters.
s% [a-zA-Z]*[-]+[a-zA-Z]*[^a-zA-Z-]?N r ) bytearrayr findallextendisalpha)r filteredwordsword last_charr
r
r filter_international_wordsB s z(CharSetProber.filter_international_wordsc C s t }d}d}tt| D ]n}| ||d }|dkr