
Python codecs extension featuring CLI tools for encoding/decoding anything

CodExt is a (Python2-3 compatible) library that extends the native codecs library (namely for adding new custom encodings and character mappings) and provides 120+ new codecs, hence its name combining CODecs EXTension. It also features a guess mode for decoding multiple layers of encoding and CLI tools for convenience.
$ pip install codext
| Want to contribute a new codec ? | Want to contribute a new macro ? |
|---|---|
| Check the documentation first Then PR your new codec | your updated version of |


Using the unbase command line tool
$ codext -i test.txt encode dna-1
GTGAGCGGGTATGTGA
$ echo -en "test" | codext encode morse
- . ... -
$ echo -en "test" | codext encode braille
⠞⠑⠎⠞
$ echo -en "test" | codext encode base100
👫👜👪👫
$ echo -en "Test string" | codext encode reverse
gnirts tseT
$ echo -en "Test string" | codext encode reverse morse
--. -. .. .-. - ... / - ... . -
$ echo -en "Test string" | codext encode reverse morse dna-2
AGTCAGTCAGTGAGAAAGTCAGTGAGAAAGTGAGTGAGAAAGTGAGTCAGTGAGAAAGTCAGAAAGTGAGTGAGTGAGAAAGTTAGAAAGTCAGAAAGTGAGTGAGTGAGAAAGTGAGAAAGTC
$ echo -en "Test string" | codext encode reverse morse dna-2 octal
101107124103101107124103101107124107101107101101101107124103101107124107101107101101101107124107101107124107101107101101101107124107101107124103101107124107101107101101101107124103101107101101101107124107101107124107101107124107101107101101101107124124101107101101101107124103101107101101101107124107101107124107101107124107101107101101101107124107101107101101101107124103
$ echo -en "AGTCAGTCAGTGAGAAAGTCAGTGAGAAAGTGAGTGAGAAAGTGAGTCAGTGAGAAAGTCAGAAAGTGAGTGAGTGAGAAAGTTAGAAAGTCAGAAAGTGAGTGAGTGAGAAAGTGAGAAAGTC" | codext -d dna-2 morse reverse
test string
$ codext add-macro my-encoding-chain gzip base63 lzma base64
$ codext list macros
example-macro, my-encoding-chain
$ echo -en "Test string" | codext encode my-encoding-chain
CQQFAF0AAIAAABuTgySPa7WaZC5Sunt6FS0ko71BdrYE8zHqg91qaqadZIR2LafUzpeYDBalvE///ug4AA==
$ codext remove-macro my-encoding-chain
$ codext list macros
example-macro
baseXX CLI tools) Playing with base encodings.
$ echo "Test string !" | base122
*.7!ft9�-f9Â
$ echo "Test string !" | base91
"ONK;WDZM%Z%xE7L
$ echo "Test string !" | base91 | base85
B2P|BJ6A+nO(j|-cttl%
$ echo "Test string !" | base91 | base85 | base36 | base58-flickr
QVx5tvgjvCAkXaMSuKoQmCnjeCV1YyyR3WErUUErFf
$ echo "Test string !" | base91 | base85 | base36 | base58-flickr | base58-flickr -d | base36 -d | base85 -d | base91 -d
Test string !
$ echo "Test string !" | base91 | base85 | base36 | base58-flickr | unbase -m 3
Test string !
$ echo "Test string !" | base91 | base85 | base36 | base58-flickr | unbase -f Test
Test string !
Listing codecs.
$ codext list encodings
a1z26 adler32 affine alternative-rot ascii
atbash autoclave bacon barbie base
base1 base2 base3 base4 base8
<<snipped>>
Finding a codec based on a name.
$ codext search bitcoin
base58
Encoding a string.
$ echo -en "This is a test" | codext encode polybius
44232443 2443 11 44154344
Encoding a file.
$ echo -en "this is a test" > to_be_encoded.txt
$ codext encode base64 < to_be_encoded.txt > text.b64
$ cat text.b64
dGhpcyBpcyBhIHRlc3Q=
Chaining codecs.
$ echo -en "mrdvm6teie6t2cq=" | codext encode upper | codext decode base32 | codext decode base64
test
Iteratively guessing decodings.
$ echo -en "test" | codext encode base64 gzip | codext guess
Codecs: gzip
dGVzdA==
$ echo -en "test" | codext encode base64 gzip | codext guess gzip -i base
Codecs: gzip, base64
test
Getting the list of available codecs.
>>> import codext
>>> codext.list()
['ascii85', 'base85', 'base100', 'base122', ..., 'tomtom', 'dna', 'html', 'markdown', 'url', 'resistor', 'sms', 'whitespace', 'whitespace-after-before']
Playing with some base encodings.
```python
>>> codext.encode("this is a test", "base58-bitcoin")
'jo91waLQA1NNeBmZKUF'
>>> codext.encode("this is a test", "base58-ripple")
'jo9rA2LQwr44eBmZK7E'
>>> codext.encode("this is a test", "base58-url")
'JN91Wzkpa1nnDbLyjtf'
>>> codecs.encode("this is a test", "base100")
'👫👟👠👪🐗👠👪🐗👘🐗👫👜👪👫'
>>> codecs.decode("👫👟👠👪🐗👠👪🐗👘🐗👫👜👪👫", "base100")
'this is a test'
Playing with some cryptography-based codecs.
>>> codext.encode("This is a test !", "vigenere-MYSECRETKET")
'Ffaw kj e mowm !'
>>> codext.encode("This is a test !", "autoclave-SECRET")
'Llkj ml t amkb !'
Encoding/decoding with various other codecs.
>>> for i in range(8):
print(codext.encode("this is a test", "dna-%d" % (i + 1)))
GTGAGCCAGCCGGTATACAAGCCGGTATACAAGCAGACAAGTGAGCGGGTATGTGA
CTCACGGACGGCCTATAGAACGGCCTATAGAACGACAGAACTCACGCCCTATCTCA
ACAGATTGATTAACGCGTGGATTAACGCGTGGATGAGTGGACAGATAAACGCACAG
AGACATTCATTAAGCGCTCCATTAAGCGCTCCATCACTCCAGACATAAAGCGAGAC
TCTGTAAGTAATTCGCGAGGTAATTCGCGAGGTAGTGAGGTCTGTATTTCGCTCTG
TGTCTAACTAATTGCGCACCTAATTGCGCACCTACTCACCTGTCTATTTGCGTGTC
GAGTGCCTGCCGGATATCTTGCCGGATATCTTGCTGTCTTGAGTGCGGGATAGAGT
CACTCGGTCGGCCATATGTTCGGCCATATGTTCGTCTGTTCACTCGCCCATACACT
>>> codext.decode("GTGAGCCAGCCGGTATACAAGCCGGTATACAAGCAGACAAGTGAGCGGGTATGTGA", "dna-1")
'this is a test'
>>> codecs.encode("this is a test", "morse")
'- .... .. ... / .. ... / .- / - . ... -'
>>> codecs.decode("- .... .. ... / .. ... / .- / - . ... -", "morse")
'this is a test'
>>> with open("morse.txt", 'w', encoding="morse") as f:
f.write("this is a test")
14
>>> with open("morse.txt",encoding="morse") as f:
f.read()
'this is a test'
>>> print(codext.encode("An example test string", "baudot-tape"))
***.**
. *
***.*
* .
.*
* .*
. *
** .*
***.**
** .**
.*
* .
* *. *
.*
* *.
* *. *
* .
* *.
* *. *
***.
*.*
***.*
* .*
base1: useless, but for the sake of completenessbase2: simple conversion to binary (with a variant with a reversed alphabet)base3: conversion to ternary (with a variant with a reversed alphabet)base4: conversion to quarternary (with a variant with a reversed alphabet)base8: simple conversion to octal (with a variant with a reversed alphabet)base10: simple conversion to decimalbase11: conversion to digits with a "a"This category also contains ascii85, adobe, [x]btoa, zeromq with the base85 codec.
baudot: supports CCITT-1, CCITT-2, EU/FR, ITA1, ITA2, MTK-2 (Python3 only), UK, ...baudot-spaced: variant of baudot ; groups of 5 bits are whitespace-separatedbaudot-tape: variant of baudot ; outputs a string that looks like a perforated tapebcd: Binary Coded Decimal, encodes characters from their (zero-left-padded) ordinalsbcd-extended0: variant of bcd ; encodes characters from their (zero-left-padded) ordinals using prefix bits 0000adler: Adler32 algorithm (relies on zlib)crc: CRC of lengths 8, 10-17, 21, 24, 30-32, 40, 64, 82 with a variety of polynomsluhn: Luhn mod N algorithma1z26: keeps words whitespace-separated and uses a custom character separatorcases: set of case-related encodings (including camel-, kebab-, lower-, pascal-, upper-, snake- and swap-case, slugify, capitalize, title)dummy: set of simple encodings (including integer, replace, reverse, word-reverse, substite and strip-spaces)octal: dummy octal conversion (converts to 3-digits groups)octal-spaced: variant of octal ; dummy octal conversion, handling whitespace separatorsordinal: dummy character ordinals conversion (converts to 3-digits groups)ordinal-spaced: variant of ; dummy character ordinals conversion, handling whitespace separatorsgzip: standard Gzip compression/decompressionlz77: compresses the given data with the algorithm of Lempel and Ziv of 1977lz78: compresses the given data with the algorithm of Lempel and Ziv of 1978pkzip_deflate: standard Zip-deflate compression/decompressionpkzip_bzip2: standard BZip2 compression/decompressionpkzip_lzma: standard LZMA compression/decompression⚠️ Compression functions are of course definitely NOT encoding functions ; they are implemented for leveraging the
.encode(...)API fromcodecs.
affine: aka Affine Cipheratbash: aka Atbash Cipherautoclave: aka Autoclave/Autokey Cipher (variant of Vigenere Cipher)bacon: aka Baconian Cipherbarbie-N: aka Barbie Typewriter (N belongs to [1, 4])beaufort: aka Beaufort Cipher (variant of Vigenere Cipher)citrix: aka Citrix CTX1 password encodingphillips: aka Phillips Cipher (polyalphabetic block cipher with 8 key squares)⚠️ Crypto functions are of course definitely NOT encoding functions ; they are implemented for leveraging the
.encode(...)API fromcodecs.
blake: includes BLAKE2b and BLAKE2s (Python 3 only ; relies on hashlib)crypt: Unix's crypt hash for passwords (Python 3 and Unix only ; relies on crypt)md: aka Message Digest ; includes MD4 and MD5 (relies on hashlib)sha: aka Secure Hash Algorithms ; includes SHA1, 224, 256, 384, 512 (Python2/3) but also SHA3-224, -256, -384 and -512 (Python 3 only ; relies on hashlib)shake: aka SHAKE hashing (Python 3 only ; relies on hashlib)⚠️ Hash functions are of course definitely NOT encoding functions ; they are implemented for convenience with the
.encode(...)API fromcodecsand useful for chaning codecs.
braille: well-known braille language (Python 3 only)ipsum: aka lorem ipsumgalactic: aka galactic alphabet or Minecraft enchantment language (Python 3 only)leetspeak: based on minimalistic elite speaking rulesmorse: uses whitespace as a separatornavajo: only handles letters (not full words from the Navajo dictionary)radio: aka NATO or radio phonetic alphabetsouthpark: converts letters to Kenny's language from Southpark (whitespace is also handled)dna: implements the 8 rules of DNA sequences (N belongs to [1,8])letter-indices: encodes consonants and/or vowels with their corresponding indicesmarkdown: unidirectional encoding from Markdown to HTMLhexagram: uses Base64 and encodes the result to a charset of I Ching hexagrams (as implemented here)klopf: aka Klopf code ; Polybius square with trivial alphabetical distributionresistor: aka resistor color codesrick: aka Rick cipher (in reference to Rick Astley's song "Never gonna give you up")sms: also called T9 code ; uses "-" as a separator for encoding, "-" or "_" or whitespace for decodinghtml: implements entities according to this referenceurl: aka URL encodingbase16base26: conversion to alphabet lettersbase32: classical conversion according to the RFC4648 with all its variants (zbase32, extended hexadecimal, geohash, Crockford)base36: Base36 conversion to letters and digits (with a variant inverting both groups)base45: Base45 DRAFT algorithm (with a variant inverting letters and digits)base58: multiple versions of Base58 (bitcoin, flickr, ripple)base62: Base62 conversion to lower- and uppercase letters and digits (with a variant with letters and digits inverted)base63: similar to base62 with the "_" addedbase64: classical conversion according to RFC4648 with its variant URL (or file) (it also holds a variant with letters and digits inverted)base67: custom conversion using some more special characters (also with a variant with letters and digits inverted)base91: Base91 custom conversionbase100 (or emoji): Base100 custom conversionbase122: Base100 custom conversionbase-genericN: see base encodings ; supports any possible basebcd-extended1bcd1111excess3: uses Excess-3 (aka Stibitz code) binary encoding to convert characters from their ordinalsgray: aka reflected binary codemanchester: XORes each bit of the input with 01manchester-inverted: variant of manchester ; XORes each bit of the input with 10rotateN: rotates characters by the specified number of bits (N belongs to [1, 7] ; Python 3 only)ordinalpolybius: aka Polybius Square Cipherrailfence: aka Rail Fence CipherrotN: aka Caesar cipher (N belongs to [1,25])scytaleN: encrypts using the number of letters on the rod (N belongs to [1,[)shiftN: shift ordinals (N belongs to [1,255])trithemius: aka Trithemius Cipher (variant of Vigenere Cipher)vigenere: aka Vigenere CipherxorN: XOR with a single byte (N belongs to [1,255])southpark-icase: case insensitive variant of southparktap: converts text to tap/knock code, commonly used by prisonerstomtom: similar to morse, using slashes and backslasheswhitespacewhitespace_after_before: variant of whitespace ; encodes characters as new characters with whitespaces before and after according to an equation described in the codec name (e.g. "whitespace+2*after-3*before")