从TikTok抓取并下载有用的信息。
这不是官方API支持等。这只是一个使用TikTok Web API来抓取媒体和相关元信息的爬虫。
### 近期活动中隐写术的最新应用
- **APT29 - 2025-03**
- APT29(Cozy Bear)被发现利用隐写术隐藏其他恶意软件阶段。他们利用公共云服务,并将加密载荷嵌入图像文件中以逃避检测。
- 参考文献:[APT29 Uses Steganography to Hide Backdoor in Image Files](https://www.trendmicro.com/vinfo/us/security/news/cybercrime-and-digital-threats/apt29-uses-steganography-to-hide-backdoor-in-image-files)
- **RedCurl - 2025-02**
- RedCurl网络犯罪团伙通过将恶意代码隐藏在图像文件中来传播恶意软件。
- 参考文献:[RedCurl Uses Steganography to Deliver Malware](https://www.bitdefender.com/en-us/blog/labs/redcurl-uses-steganography-to-deliver-malware)
- **DLSymbiote - 2024-09**
- 发现了一种名为DLSymbiote的新型Linux恶意软件,它使用隐写术将载荷隐藏在图像中。
- 参考文献:[DLSymbiote Uses Steganography to Hide Payloads](https://www.welivesecurity.com/en/eset-research/dlsymbiote-linux-malware-uses-steganography/)
- **Mallox - 2024-08**
- Mallox勒索软件团伙被发现使用隐写术将恶意代码隐藏在图像文件中。
- 参考文献:[Mallox Ransomware Uses Steganography](https://www.bleepingcomputer.com/news/security/mallox-ransomware-uses-steganography-to-hide-malicious-code-in-images/)
- **SteganoAmor - 2024-06**
- TA558黑客组织被发现使用隐写术将恶意代码隐藏在图像中,并将其发送到目标系统。
- 参考文献:[TA558 Uses Steganography to Deliver Malware](https://www.bleepingcomputer.com/news/security/ta558-hackers-use-steganography-to-deliver-malware-in-images/)
- **Go-Go-Stego - 2024-05**
- 发现了一场针对印度、美国及其他国家用户的新型隐写术活动。攻击链涉及使用嵌入了恶意代码的图像文件。
- 参考文献:[Go-Go-Stego Campaign](https://www.forcepoint.com/blog/x-labs/go-go-stego-campaign-india-us)
- **SteganoBalto - 2024-03**
- 揭露了一场活跃的隐写术活动,针对拉丁美洲(LATAM)用户。攻击者利用隐写术将恶意载荷隐藏在看似无害的图像文件中。
- 参考文献:[SteganoBalto Targets Latin America with Steganography](https://www.forcepoint.com/blog/x-labs/steganobalto-targets-latin-america)
- **Agent Tesla - 2024-02**
- Agent Tesla键盘记录器被发现使用隐写术将载荷隐藏在图像中。
- 参考文献:[Agent Tesla Uses Steganography to Hide Payloads](https://www.bleepingcomputer.com/news/security/agent-tesla-malware-uses-steganography-to-hide-payloads-in-images/)
- **SteganoAmor – TA558 - 2024-02**
- TA558黑客组织被发现使用隐写术将恶意代码隐藏在图像中,并将其发送到目标系统。
- 参考文献:[TA558 Uses Steganography to Deliver Malware](https://www.bleepingcomputer.com/news/security/ta558-hackers-use-steganography-to-deliver-malware-in-images/)
- **Mystic Stealer - 2023-12**
- Mystic Stealer恶意软件被观察到使用隐写术隐藏其载荷。
- 参考文献:[Mystic Stealer Uses Steganography](https://www.bleepingcomputer.com/news/security/mystic-stealer-malware-uses-steganography-to-hide-payloads-in-images/)
- **Gamaredon - 2023-11**
- Gamaredon APT组织被看到使用隐写术将恶意代码隐藏在图像文件中。
- 参考文献:[Gamaredon Uses Steganography](https://www.bleepingcomputer.com/news/security/gamaredon-apt-group-uses-steganography-to-hide-malicious-code-in-images/)
- **Steganography in Malware - 2023-10**
- 针对各种恶意软件家族如何利用隐写术在图像中隐藏载荷的详细分析。
- 参考文献:[Steganography in Malware Analysis](https://www.mandiant.com/resources/blog/steganography-malware-analysis)
- **Agent Tesla - 2023-09**
- Agent Tesla键盘记录器被发现使用隐写术将载荷隐藏在图像中。
- 参考文献:[Agent Tesla Uses Steganography](https://www.bleepingcomputer.com/news/security/agent-tesla-malware-uses-steganography-to-hide-payloads-in-images/)```sh
yarn build
tiktok-scraper 需要 Node.js v10+ 才能运行。
从 NPM 安装```sh npm i -g tiktok-scraper
**从 YARN 安装**```sh
yarn global add tiktok-scraper
$ tiktok-scraper --help
Usage: tiktok-scraper [options]
Commands: tiktok-scraper user [id] Scrape videos from username. Enter only username tiktok-scraper hashtag [id] Scrape videos from hashtag. Enter hashtag without # tiktok-scraper trend Scrape posts from current trends tiktok-scraper music [id] Scrape posts from a music id number tiktok-scraper video [id] Download single video without the watermark tiktok-scraper history View previous download history tiktok-scraper from-file [file] [async] Scrape users, hashtags, music, videos mentioned in a file. 1 value per 1 line
Options: --version Show version number [boolean] --session Set session cookie value. Sometimes session can be helpful when scraping data from any method [default: ""] --session-file Set path to the file with list of active sessions. One session per line! [default: ""] --timeout Set timeout between requests. Timeout is in Milliseconds: 1000 mls = 1 s [default: 0] --number, -n Number of posts to scrape. If you will set 0 then all posts will be scraped [default: 0] --since Scrape no posts published before this date (timestamp). If set to 0 the filter is deactived [default: 0] --proxy, -p Set single proxy [default: ""] --proxy-file Use proxies from a file. Scraper will use random proxies from the file per each request. 1 line 1 proxy. [default: ""] --download, -d Download video posts to the folder with the name input [id] [boolean] [default: false] --asyncDownload, -a Number of concurrent downloads [default: 5] --hd Download video in HD. Video size will be x5-x10 times larger and this will affect scraper execution speed. This option only works in combination with -w flag [boolean] [default: false] --zip, -z ZIP all downloaded video posts [boolean] [default: false] --filepath File path to save all output files. [default: "/Users/karl.wint/Documents/projects/javascript/tiktok-scraper"] --filetype, -t Type of the output file where post information will be saved. 'all' - save information about all posts to the` 'json' and 'csv' [choices: "csv", "json", "all", ""] [default: ""] --filename, -f Set custom filename for the output files [default: ""] --noWaterMark, -w Download video without the watermark. NOTE: With the recent update you only need to use this option if you are scraping Hashtag Feed. User/Trend/Music feeds will have this url by default [boolean] [default: false] --store, -s Scraper will save the progress in the OS TMP or Custom folder and in the future usage will only download new videos avoiding duplicates [boolean] [default: false] --historypath Set custom path where history file/files will be stored [default: "/var/folders/d5/fyh1_f2926q7c65g7skc0qh80000gn/T"] --remove, -r Delete the history record by entering "TYPE:INPUT" or "all" to clean all the history. For example: user:bob [default: ""] --webHookUrl Set webhook url to receive scraper result as HTTP requests. For example to your own API [default: ""] --method Receive data to your webhook url as POST or GET request [choices: "GET", "POST"] [default: "POST"] --help Show help [boolean]
Examples: tiktok-scraper user USERNAME -d -n 100 --session sid_tt=dae32131231 tiktok-scraper trend -d -n 100 --session sid_tt=dae32131231 tiktok-scraper hashtag HASHTAG_NAME -d -n 100 --session sid_tt=dae32131231 tiktok-scraper music MUSIC_ID -d -n 50 --session sid_tt=dae32131231 tiktok-scraper video https://www.tiktok.com/@tiktok/video/6807491984882765062 -d tiktok-scraper history tiktok-scraper history -r user:bob tiktok-scraper history -r all tiktok-scraper from-file BATCH_FILE ASYNC_TASKS -d
- [终端示例](https://github.com/drawrowfly/tiktok-scraper/tree/master/examples/CLI/Examples.md)
- [管理下载历史](https://github.com/drawrowfly/tiktok-scraper/tree/master/examples/CLI/DownloadHistory.md)
- [批量抓取与下载](https://github.com/drawrowfly/tiktok-scraper/tree/master/examples/CLI/BatchDownload.md)
### 输出文件示例

## Docker
使用 Docker 时,您将无法使用 --filepath 和 --historypath,但您可以使用 -v 设置卷(**保存所有文件的主机路径**)
##### 构建```sh
docker build . -t tiktok-scraper
示例1: 包括历史文件在内的所有文件都将保存到你运行 Docker 的目录($pwd)中。```sh docker run -v $(pwd):/usr/app/files tiktok-scraper user tiktok -d -n 5 -s
**示例2:** 包括历史文件在内的所有文件将保存在 /User/blah/downloads```sh
docker run -v /User/blah/downloads:/usr/app/files tiktok-scraper user tiktok -d -n 5 -s
.user(id, options) //Scrape posts from a specific user (Promise) .hashtag(id, options) //Scrape posts from hashtag section (Promise) .trend('', options) // Scrape posts from a trends section (Promise) .music(id, options) // Scrape posts by music id (Promise)
.userEvent(id, options) //Scrape posts from a specific user (Event) .hashtagEvent(id, options) //Scrape posts from hashtag section (Event) .trendEvent('', options) // Scrape posts from a trends section (Event) .musicEvent(id, options) // Scrape posts by music id (Event)
.getUserProfileInfo('USERNAME', options) // Get user profile information .getHashtagInfo('HASHTAG', options) // Get hashtag information .signUrl('URL', options) // Get signature for the request .getVideoMeta('WEB_VIDEO_URL', options) // Get video meta info, including video url without the watermark .getMusicInfo('https://www.tiktok.com/music/original-sound-6801885499343571718', options) // Get music metadata
### 选项```javascript
const options = {
// Number of posts to scrape: {int default: 20}
number: 50,
// Scrape posts published since this date: { int default: 0}
since: 0,
// Set session: {string[] default: ['']}
// Authenticated session cookie value is required to scrape user/trending/music/hashtag feed
// You can put here any number of sessions, each request will select random session from the list
sessionList: ['sid_tt=21312213'],
// Set proxy {string[] | string default: ''}
// http proxy: 127.0.0.1:8080
// socks proxy: socks5://127.0.0.1:8080
// You can pass proxies as an array and scraper will randomly select a proxy from the array to execute the requests
proxy: '',
// Set to {true} to search by user id: {boolean default: false}
by_user_id: false,
// How many post should be downloaded asynchronously. Only if {download:true}: {int default: 5}
asyncDownload: 5,
// How many post should be scraped asynchronously: {int default: 3}
// Current option will be applied only with current types: music and hashtag
// With other types it is always 1 because every request response to the TikTok API is providing the "maxCursor" value
// that is required to send the next request
asyncScraping: 3,
// File path where all files will be saved: {string default: 'CURRENT_DIR'}
filepath: `CURRENT_DIR`,
// Custom file name for the output files: {string default: ''}
fileName: `CURRENT_DIR`,
// Output with information can be saved to a CSV or JSON files: {string default: 'na'}
// 'csv' to save in csv
// 'json' to save in json
// 'all' to save in json and csv
// 'na' to skip this step
filetype: `na`,
// Set custom headers: user-agent, cookie and etc
// NOTE: When you parse video feed or single video metadata then in return you will receive {headers} object
// that was used to extract the information and in order to access and download video through received {videoUrl} value you need to use same headers
headers: {
'user-agent': "BLAH",
referer: 'https://www.tiktok.com/',
cookie: `tt_webid_v2=68dssds`,
},
// Download video without the watermark: {boolean default: false}
// Set to true to download without the watermark
// This option will affect the execution speed
noWaterMark: false,
// Create link to HD video: {boolean default: false}
// This option will only work if {noWaterMark} is set to {true}
hdVideo: false,
// verifyFp is used to verify the request and avoid captcha
// When you are using proxy then there are high chances that the request will be
// blocked with captcha
// You can set your own verifyFp value or default(hardcoded) will be used
verifyFp: '',
// Switch main host to Tiktok test enpoint.
// When your requests are blocked by captcha you can try to use Tiktok test endpoints.
useTestEndpoints: false
};
别忘了查看 examples 文件夹
const TikTokScraper = require('tiktok-scraper');
// User feed by username (async () => { try { const posts = await TikTokScraper.user('USERNAME', { number: 100, sessionList: ['sid_tt=58ba9e34431774703d3c34e60d584475;'] }); console.log(posts); } catch (error) { console.log(error); } })();
// User feed by user id
// Some TikTok user id's are larger then MAX_SAFE_INTEGER, you need to pass user id as a string
(async () => {
try {
const posts = await TikTokScraper.user(USER_ID, {
number: 100,
by_user_id: true,
sessionList: ['sid_tt=58ba9e34431774703d3c34e60d584475;']
});
console.log(posts);
} catch (error) {
console.log(error);
}
})();
// Trending feed (async () => { try { const posts = await TikTokScraper.trend('', { number: 100, sessionList: ['sid_tt=58ba9e34431774703d3c34e60d584475;'] }); console.log(posts); } catch (error) { console.log(error); } })();
// Hashtag feed (async () => { try { const posts = await TikTokScraper.hashtag('HASHTAG', { number: 100, sessionList: ['sid_tt=58ba9e34431774703d3c34e60d584475;'] }); console.log(posts); } catch (error) { console.log(error); } })();
// Get single user profile information: Number of followers and etc // input - USERNAME // options - not required (async () => { try { const user = await TikTokScraper.getUserProfileInfo('USERNAME', options); console.log(user); } catch (error) { console.log(error); } })();
// Get single hashtag information: Number of views and etc // input - HASHTAG NAME // options - not required (async () => { try { const hashtag = await TikTokScraper.getHashtagInfo('HASHTAG', options); console.log(hashtag); } catch (error) { console.log(error); } })();
// Get single video metadata // input - WEB_VIDEO_URL // For example: https://www.tiktok.com/@tiktok/video/6807491984882765062 // options - not required (async () => { try { const videoMeta = await TikTokScraper.getVideoMeta('https://www.tiktok.com/@tiktok/video/6807491984882765062', options); console.log(videoMeta); } catch (error) { console.log(error); } })();
### 事件```javascript
const TikTokScraper = require('tiktok-scraper');
const users = TikTokScraper.userEvent("tiktok", { number: 30 });
users.on('data', json => {
//data in JSON format
});
users.on('done', () => {
//completed
});
users.on('error', error => {
//error message
});
users.scrape();
const hashtag = TikTokScraper.hashtagEvent("summer", { number: 250, proxy: 'socks5://1.1.1.1:90' });
hashtag.on('data', json => {
//data in JSON format
});
hashtag.on('done', () => {
//completed
});
hashtag.on('error', error => {
//error message
});
hashtag.scrape();
非必需
非常常见的问题是,TikTok 将您的 IP/代理列入黑名单。在这种情况下,您可以尝试设置会话,这将增加成功的几率
获取会话:
设置会话:
CLI:
sid_tt=521kkadkasdaskdj4j213j12j312;
sid_tt=521kkadkasdaskdj4j213j12j312;
sid_tt=521kkadkasdaskdj4j213j12j312;
sid_tt=521kkadkasdaskdj4j213j12j312;
在 MODULE 中,您可以通过设置选项值 sessionList 来设置会话。例如 sessionList:["sid_tt=521kkadkasdaskdj4j213j12j312;", "sid_tt=12312312312312;"]
此部分与 MODULE 使用相关(而非 CLI)
{videoUrl} 值与 cookie 值 {tt_webid_v2} 绑定,后者可以包含 任意值
当您从用户、标签、音乐、热门推荐或单个视频中提取视频时,在响应中除了视频元数据外,您还会收到 headers 对象,其中包含用于提取数据的参数。重要部分在于,要通过 {videoUrl} 值访问/下载视频,您需要使用相同的 {headers} 值。```json headers: { "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.80 Safari/537.36", "referer": "https://www.tiktok.com/", "cookie": "tt_webid_v2=689854141086886123" },
#### 方法 2:自定义标头
您可以使用 **{options}** 传递自己的标头。```javascript
const headers = {
"user-agent": "BOB",
"referer": "https://www.tiktok.com/",
"cookie": "tt_webid_v2=BOB"
}
getVideoMeta('WEB_VIDEO_URL', {headers})
user('WEB_VIDEO_URL', {headers})
hashtag('WEB_VIDEO_URL', {headers})
trend('WEB_VIDEO_URL', {headers})
music('WEB_VIDEO_URL', {headers})
// And after you can access video through {videoUrl} value by using same custom headers
方法输出示例:user, hashtag, trend, music, userEvent, hashtagEvent, musicEvent, trendEvent```javascript { headers: { 'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.80 Safari/537.36', referer: 'https://www.tiktok.com/', cookie: 'tt_webid_v2=689854141086886123' }, collector:[{ id: 'VIDEO_ID', text: 'CAPTION', createTime: '1583870600', authorMeta:{ id: 'USER ID', name: 'USERNAME', following: 195, fans: 43500, heart: '1093998', video: 3, digg: 95, verified: false, private: false, signature: 'USER BIO', avatar:'AVATAR_URL' }, musicMeta:{ musicId: '6808098113188120838', musicName: 'blah blah', musicAuthor: 'blah', musicOriginal: true, playUrl: 'SOUND/MUSIC_URL', }, covers:{ default: 'COVER_URL', origin: 'COVER_URL', dynamic: 'COVER_URL' }, imageUrl:'IMAGE_URL', videoUrl:'VIDEO_URL', videoUrlNoWaterMark:'VIDEO_URL_WITHOUT_THE_WATERMARK', videoMeta: { width: 480, height: 864, ratio: 14, duration: 14 }, diggCount: 2104, shareCount: 1, playCount: 9007, commentCount: 50, mentions: ['@bob', '@sam', '@bob_again', '@and_sam_again'], hashtags: [{ id: '69573911', name: 'PlayWithLife', title: 'HASHTAG_TITLE', cover: [Array] }...], downloaded: true }...], //If {filetype} and {download} options are enbabled then: zip: '/{CURRENT_PATH}/user_1552963581094.zip', json: '/{CURRENT_PATH}/user_1552963581094.json', csv: '/{CURRENT_PATH}/user_1552963581094.csv' }
##### getUserProfileInfo```javascript
{
secUid: 'MS4wLjABAAAAv7iSuuXDJGDvJkmH_vz1qkDZYo1apxgzaxdBSeIuPiM',
userId: '107955',
isSecret: false,
uniqueId: 'tiktok',
nickName: 'TikTok',
signature: 'Make Your Day',
covers: ['COVER_URL'],
coversMedium: ['COVER_URL'],
following: 490,
fans: 38040567,
heart: '211522962',
video: 93,
verified: true,
digg: 29,
}
{ challengeId: '4231', challengeName: 'love', text: '', covers: [], coversMedium: [], posts: 66904972, views: '194557706433', isCommerce: false, splitTitle: '' }
##### getVideoMeta```javascript
{
headers: {
'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.80 Safari/537.36',
referer: 'https://www.tiktok.com/',
cookie: 'tt_webid_v2=689854141086886123'
},
collector:[{
id: '6807491984882765062',
text: 'We’re kicking off the #happyathome live stream series today at 5pm PT!',
createTime: '1584992742',
authorMeta: { id: '6812221792183403526', name: 'blah' },
musicMeta:{
musicId: '6822233276137213677',
musicName: 'blah',
musicAuthor: 'blah'
},
imageUrl: 'IMAGE_URL',
videoUrl: 'VIDEO_URL',
videoUrlNoWaterMark: 'VIDEO_URL_WITHOUT_THE_WATERMARK',
videoMeta: { width: 480, height: 864, ratio: 14, duration: 14 },
covers:{
default: 'COVER_URL',
origin: 'COVER_URL'
},
diggCount: 49292,
shareCount: 339,
playCount: 614678,
commentCount: 4023,
downloaded: false,
hashtags: [],
}]
}
{ music: { id: '6882925279036066566', title: 'doja x calabria', playUrl: 'dfdfdfdf', coverThumb: 'dfdfdf', coverMedium: 'dfdfdf', coverLarge: 'fdfdf', authorName: 'bryce', original: true, playToken: 'ffdfdf', keyToken: 'dfdfdfd', audioURLWithcookie: false, private: false, duration: 46, album: '', }, author: { id: '6835300004094166021', uniqueId: 'mashupsbybryce', nickname: 'bryce', avatarThumb: 'dfdfd', avatarMedium: 'dfdfdf', avatarLarger: 'dfdfdf', signature: 'hi ily :)\n70k sounds cool tbh\n👇follow my soundcloud & insta👇', verified: false, secUid: 'MS4wLjABAAAA1_5bjLAamayD4rv3q49qJGa_7dZ5jzExTO0ozOybqIwwhw5TAg_iM25lkO94DM3K', secret: false, ftc: false, relation: 0, openFavorite: false, commentSetting: 0, duetSetting: 0, stitchSetting: 0, privateAccount: false, }, stats: { videoCount: 361700 }, shareMeta: { title: 'bryceyouloser | ♬ doja x calabria | on TikTok', desc: '361.0k videos - Watch awesome short ' + 'videos created with ♬ doja x calabria', }, };
<a href="https://www.buymeacoffee.com/Usom2qC" target="_blank"><img src="https://assets.kitploit.com/production/public/readmes/5141/c20d7d7728f977412007ddec6509758f76c0727e972aea2c4c18ead6c69ff2af.png" alt="Buy Me A Coffee" style="height: 41px !important;width: 174px !important;box-shadow: 0px 3px 2px 0px rgba(190, 190, 190, 0.5) !important;-webkit-box-shadow: 0px 3px 2px 0px rgba(190, 190, 190, 0.5) !important;" ></a>
---
许可证
---
**MIT**
**自由软件**