/tmp filled with bszip-###### #5090

Closed
opened 2026-02-05 09:39:19 +03:00 by OVERLORD · 5 comments
Owner

Originally created by @JtheBAB on GitHub (Dec 28, 2024).

Describe the Bug

I upgraded a few days before to 24.12 and now my /tmp is full with a lot of files like:

-rw------- 1 www-data www-data 131704 Dec 26 19:56 bszip-zjPTTH
-rw------- 1 www-data www-data 11347585 Dec 27 19:56 bszip-ZlGVV9
-rw------- 1 www-data www-data 1797 Dec 26 11:56 bszip-zmQv5W
-rw------- 1 www-data www-data 1356 Dec 26 03:56 bszip-ZnG94v
-rw------- 1 www-data www-data 78599 Dec 26 03:56 bszip-zo1YiQ
-rw------- 1 www-data www-data 719 Dec 26 11:56 bszip-zo6xRj
-rw------- 1 www-data www-data 25234 Dec 26 19:56 bszip-ZobfbT
-rw------- 1 www-data www-data 25236 Dec 26 23:56 bszip-ZQFTFx
-rw------- 1 www-data www-data 19664 Dec 28 03:57 bszip-zrvqwX
-rw------- 1 www-data www-data 19666 Dec 27 03:56 bszip-zSwzym
-rw------- 1 www-data www-data 1160 Dec 27 07:56 bszip-zt8YxX
-rw------- 1 www-data www-data 31438858 Dec 27 23:56 bszip-zufGgj
-rw------- 1 www-data www-data 0 Dec 28 11:56 bszip-zYAjYr
-rw------- 1 www-data www-data 1372 Dec 25 23:57 bszip-zZAkKm
-rw------- 1 www-data www-data 7707551 Dec 27 15:56 bszip-zzex3O
-rw------- 1 www-data www-data 381 Dec 26 23:56 bszip-ZZnwjc
-rw------- 1 www-data www-data 802216 Dec 25 23:57 bszip-ZzRR3m

When i unzip it than i can see that is random content from my bookstack instance.

I also see this:

-rw------- 1 www-data www-data 397305 Dec 21 19:55 bs-pdfgen-html-00j56x
-rw------- 1 www-data www-data 22275 Dec 21 03:56 bs-pdfgen-html-00Rg25
-rw------- 1 www-data www-data 55713 Dec 19 19:56 bs-pdfgen-html-01hIYc
-rw------- 1 www-data www-data 132196 Dec 27 11:56 bs-pdfgen-html-023StT
-rw------- 1 www-data www-data 397305 Dec 23 03:55 bs-pdfgen-html-04EyQT
-rw------- 1 www-data www-data 0 Dec 28 11:56 bs-pdfgen-html-06Md6w
-rw------- 1 www-data www-data 0 Dec 28 11:56 bs-pdfgen-html-06P4IX
-rw------- 1 www-data www-data 8529630 Dec 26 07:55 bs-pdfgen-html-0c6wnA
-rw------- 1 www-data www-data 20693 Dec 24 23:56 bs-pdfgen-html-0deTtg
-rw------- 1 www-data www-data 22919 Dec 21 03:56 bs-pdfgen-html-0DqgSL
-rw------- 1 www-data www-data 397305 Dec 23 23:56 bs-pdfgen-html-0fOOfE
-rw------- 1 www-data www-data 90275 Dec 21 11:55 bs-pdfgen-html-0fs1cD

I really doubt that someone export random pages from Bookstack (at least nothing in the audit log - if logged at all). Any idea what that could be?

Steps to Reproduce

I just updated and then a few days later i got the message from my monitoring system that the /tmp is full

Expected Behaviour

It should only create a file when needed. Cleanup old files?

Screenshots or Additional Context

No response

Browser Details

No response

Exact BookStack Version

v24.12

Originally created by @JtheBAB on GitHub (Dec 28, 2024). ### Describe the Bug I upgraded a few days before to 24.12 and now my /tmp is full with a lot of files like: -rw------- 1 www-data www-data 131704 Dec 26 19:56 bszip-zjPTTH -rw------- 1 www-data www-data 11347585 Dec 27 19:56 bszip-ZlGVV9 -rw------- 1 www-data www-data 1797 Dec 26 11:56 bszip-zmQv5W -rw------- 1 www-data www-data 1356 Dec 26 03:56 bszip-ZnG94v -rw------- 1 www-data www-data 78599 Dec 26 03:56 bszip-zo1YiQ -rw------- 1 www-data www-data 719 Dec 26 11:56 bszip-zo6xRj -rw------- 1 www-data www-data 25234 Dec 26 19:56 bszip-ZobfbT -rw------- 1 www-data www-data 25236 Dec 26 23:56 bszip-ZQFTFx -rw------- 1 www-data www-data 19664 Dec 28 03:57 bszip-zrvqwX -rw------- 1 www-data www-data 19666 Dec 27 03:56 bszip-zSwzym -rw------- 1 www-data www-data 1160 Dec 27 07:56 bszip-zt8YxX -rw------- 1 www-data www-data 31438858 Dec 27 23:56 bszip-zufGgj -rw------- 1 www-data www-data 0 Dec 28 11:56 bszip-zYAjYr -rw------- 1 www-data www-data 1372 Dec 25 23:57 bszip-zZAkKm -rw------- 1 www-data www-data 7707551 Dec 27 15:56 bszip-zzex3O -rw------- 1 www-data www-data 381 Dec 26 23:56 bszip-ZZnwjc -rw------- 1 www-data www-data 802216 Dec 25 23:57 bszip-ZzRR3m When i unzip it than i can see that is random content from my bookstack instance. I also see this: -rw------- 1 www-data www-data 397305 Dec 21 19:55 bs-pdfgen-html-00j56x -rw------- 1 www-data www-data 22275 Dec 21 03:56 bs-pdfgen-html-00Rg25 -rw------- 1 www-data www-data 55713 Dec 19 19:56 bs-pdfgen-html-01hIYc -rw------- 1 www-data www-data 132196 Dec 27 11:56 bs-pdfgen-html-023StT -rw------- 1 www-data www-data 397305 Dec 23 03:55 bs-pdfgen-html-04EyQT -rw------- 1 www-data www-data 0 Dec 28 11:56 bs-pdfgen-html-06Md6w -rw------- 1 www-data www-data 0 Dec 28 11:56 bs-pdfgen-html-06P4IX -rw------- 1 www-data www-data 8529630 Dec 26 07:55 bs-pdfgen-html-0c6wnA -rw------- 1 www-data www-data 20693 Dec 24 23:56 bs-pdfgen-html-0deTtg -rw------- 1 www-data www-data 22919 Dec 21 03:56 bs-pdfgen-html-0DqgSL -rw------- 1 www-data www-data 397305 Dec 23 23:56 bs-pdfgen-html-0fOOfE -rw------- 1 www-data www-data 90275 Dec 21 11:55 bs-pdfgen-html-0fs1cD I really doubt that someone export random pages from Bookstack (at least nothing in the audit log - if logged at all). Any idea what that could be? ### Steps to Reproduce I just updated and then a few days later i got the message from my monitoring system that the /tmp is full ### Expected Behaviour It should only create a file when needed. Cleanup old files? ### Screenshots or Additional Context _No response_ ### Browser Details _No response_ ### Exact BookStack Version v24.12
OVERLORD added the 🐛 Bug label 2026-02-05 09:39:19 +03:00
Author
Owner

@ssddanbrown commented on GitHub (Dec 30, 2024):

Hi @JtheBAB,

I've probably been a bit lazy in regard to ensuring temp files are cleaned (assuming they'd be cleaned up by the system) but we should improve that for scenarios where they aren't cleaned as often.

Is your instance publicly accessible?
Exports are not logged in the audit log (same as other read-only activity) but access to exports can be controlled via role permissions.

@ssddanbrown commented on GitHub (Dec 30, 2024): Hi @JtheBAB, I've probably been a bit lazy in regard to ensuring temp files are cleaned (assuming they'd be cleaned up by the system) but we should improve that for scenarios where they aren't cleaned as often. Is your instance publicly accessible? Exports are not logged in the audit log (same as other read-only activity) but access to exports can be controlled via role permissions.
Author
Owner

@JtheBAB commented on GitHub (Dec 30, 2024):

Some books are available without authentication but not public like internet access. What i find strange that i have regular the same exports during the same time of a day. Like something would trigger that export.

I removed now the export functionality for the Public user.

@JtheBAB commented on GitHub (Dec 30, 2024): Some books are available without authentication but not public like internet access. What i find strange that i have regular the same exports during the same time of a day. Like something would trigger that export. I removed now the export functionality for the Public user.
Author
Owner

@ssddanbrown commented on GitHub (Dec 31, 2024):

Strange, if needed you might be able to use webserver access logs to track back who was requesting these exports.

I've assigned this for the next patch, and started work in #5379 where I also plan to also add better cleanup for PDF exports.

@ssddanbrown commented on GitHub (Dec 31, 2024): Strange, if needed you might be able to use webserver access logs to track back who was requesting these exports. I've assigned this for the next patch, and started work in #5379 where I also plan to also add better cleanup for PDF exports.
Author
Owner

@JtheBAB commented on GitHub (Jan 2, 2025):

I found now what is responsible. We have a internal Sharepoint instances that crawls also the bookstack instance. And it looks like that can trigger the export:

XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:13 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 7771476 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)"
XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 54202 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)"
XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 81087 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)"
XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 3892 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT;
MS Search 6.0 Robot)"

Don't now if "just" crawling should trigger the export.

@JtheBAB commented on GitHub (Jan 2, 2025): I found now what is responsible. We have a internal Sharepoint instances that crawls also the bookstack instance. And it looks like that can trigger the export: XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:13 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 7771476 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)" XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 54202 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)" XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 81087 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)" XXX.XXX.XXX.XX - - [29/Dec/2024:23:57:15 +0100] "GET /books/#####/######/export/zip HTTP/1.1" 200 3892 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT; MS Search 6.0 Robot)" Don't now if "just" crawling should trigger the export.
Author
Owner

@ssddanbrown commented on GitHub (Jan 5, 2025):

Don't now if "just" crawling should trigger the export.

Crawling is just fetching as a user would, so we'd have to actively be defensive (use other HTTP methods) to avoid that.
Alternatively the robots.txt could maybe be customized to avoid the bot fetching export URLs, if it listens to robots.txt files.

Either way, I've now added more substantial cleanup for exports in #5379 which will be part of the next patch release, so I'll therefore close this off.

@ssddanbrown commented on GitHub (Jan 5, 2025): > Don't now if "just" crawling should trigger the export. Crawling is just fetching as a user would, so we'd have to actively be defensive (use other HTTP methods) to avoid that. Alternatively the robots.txt could maybe be customized to avoid the bot fetching export URLs, if it listens to robots.txt files. Either way, I've now added more substantial cleanup for exports in #5379 which will be part of the next patch release, so I'll therefore close this off.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: starred/BookStack#5090