2bi,lSSKrSSKrSSKJrJr SSKJrJr \R"S5r "SS5r g)N)OptionalUnion)LanguageFilter ProbingStates%[a-zA-Z]*[-]+[a-zA-Z]*[^a-zA-Z-]?c>\rSrSrSr\R 4S\SS4SjjrSSjr\ S\ \ 4Sj5r \ S\ \ 4S j5r S \\\4S\4S jr\ S\4S j5rS\4S jr\S\\\4S\4Sj5r\S\\\4S\4Sj5r\S\\\4S\4Sj5rSrg) CharSetProber(gffffff? lang_filterreturnNc[RUlSUlXl[ R "[5Ulg)NT) r DETECTING_stateactiver logging getLogger__name__logger)selfr s ړ/builddir/build/BUILDROOT/alt-python313-pip-23.3.1-3.el8.x86_64/opt/alt/python313/lib/python3.13/site-packages/pip/_vendor/chardet/charsetprober.py__init__CharSetProber.__init__,s.",,  &''1 c.[RUlgN)rrrrs rresetCharSetProber.reset2s",, rcgrrs r charset_nameCharSetProber.charset_name5src[erNotImplementedErrorrs rlanguageCharSetProber.language9s!!rbyte_strc[err$)rr(s rfeedCharSetProber.feed=s!!rcUR$r)rrs rstateCharSetProber.state@s {{rcg)Ngr rs rget_confidenceCharSetProber.get_confidenceDsrbufc6[R"SSU5nU$)Ns([-])+ )resub)r2s rfilter_high_byte_only#CharSetProber.filter_high_byte_onlyGsff&c2 rc[5n[RU5nUHJnURUSS5 USSnUR 5(dUS:aSnURU5 ML U$)u We define three types of bytes: alphabet: english alphabets [a-zA-Z] international: international characters [€-ÿ] marker: everything else [^a-zA-Z€-ÿ] The input buffer can be thought to contain a series of words delimited by markers. This function works to filter all words that contain at least one international character. All contiguous sequences of markers are replaced by a single space ascii character. This filter applies to all scripts which do not use English characters. Nr4) bytearrayINTERNATIONAL_WORDS_PATTERNfindallextendisalpha)r2filteredwordsword last_chars rfilter_international_words(CharSetProber.filter_international_wordsLst; ,33C8D OOD"I & RS I$$&&9w+> OOI &rcD[5nSnSn[U5RS5n[U5HNupEUS:Xa US-nSnMUS:XdMXC:a+U(d$UR XU5 UR S5 SnMP U(dUR XS 5 U$) a+ Returns a copy of ``buf`` that retains only the sequences of English alphabet and high byte characters that are not between <> characters. This filter can be applied to all scripts which contain both English characters and extended ASCII characters, but is currently only used by ``Latin1Prober``. Frc>r(?IB$U5)#34$$$rr ) rr5typingrrenumsrrcompiler=r r rrrbs1: "/ jj8 kkr