Jobs are Frequently Hung #6019

Closed
opened 2026-02-05 11:56:44 +03:00 by OVERLORD · 5 comments
Owner

Originally created by @davejhahn on GitHub (May 17, 2025).

I have searched the existing issues, both open and closed, to make sure this is not a duplicate report.

  • Yes

The bug

I have a large collection, approximately 125K images, so maybe this is something that surfaces in this scenario, but I frequently have jobs that just get stuck.

e.g.
Extract Metadata, it will process for a while, I see the counts going down, and then it just hangs, as do any other jobs.

The only way I can get past it is to stop the service (docker compose down) and restart. Sometimes it works then, other times it repeats.

I've streamed the log and do not see anything being logged that indicates what might be going on.

Nothing appears in logs, at least the way I am looking at them:

docker logs --follow immich_server

The OS that Immich Server is running on

Ubuntu 22.04

Version of Immich Server

v1.32.3

Version of Immich Mobile App

1.132.3.build.205

Platform with the issue

  • Server
  • Web
  • Mobile

Your docker-compose.yml content

#
# WARNING: To install Immich, follow our guide: https://immich.app/docs/install/docker-compose
#
# Make sure to use the docker-compose.yml of the current release:
#
# https://github.com/immich-app/immich/releases/latest/download/docker-compose.yml
#
# The compose file on main may not be compatible with the latest release.

name: immich

services:
  immich-server:
    container_name: immich_server
    image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release}
    # extends:
    #   file: hwaccel.transcoding.yml
    #   service: cpu # set to one of [nvenc, quicksync, rkmpp, vaapi, vaapi-wsl] for accelerated transcoding
    volumes:
      # Do not edit the next line. If you want to change the media storage location on your system, edit the value of UPLOAD_LOCATION in the .env file
      - ${UPLOAD_LOCATION}:/usr/src/app/upload
      - /etc/localtime:/etc/localtime:ro
      - /images:/images
    env_file:
      - .env
    ports:
      - '2283:2283'
    depends_on:
      - redis
      - database
    restart: always
    healthcheck:
      disable: false

  immich-machine-learning:
    container_name: immich_machine_learning
    # For hardware acceleration, add one of -[armnn, cuda, rocm, openvino, rknn] to the image tag.
    # Example tag: ${IMMICH_VERSION:-release}-cuda
    image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}
    # extends: # uncomment this section for hardware acceleration - see https://immich.app/docs/features/ml-hardware-acceleration
    #   file: hwaccel.ml.yml
    #   service: cpu # set to one of [armnn, cuda, rocm, openvino, openvino-wsl, rknn] for accelerated inference - use the `-wsl` version for WSL2 where applica
ble
    volumes:
      - model-cache:/cache
    env_file:
      - .env
    restart: always
    healthcheck:
      disable: false

  redis:
    container_name: immich_redis
    image: docker.io/valkey/valkey:8-bookworm@sha256:42cba146593a5ea9a622002c1b7cba5da7be248650cbb64ecb9c6c33d29794b1
    healthcheck:
      test: redis-cli ping || exit 1
    restart: always

  database:
    container_name: immich_postgres
    image: docker.io/tensorchord/pgvecto-rs:pg14-v0.2.0@sha256:739cdd626151ff1f796dc95a6591b55a714f341c737e27f045019ceabf8e8c52
    environment:
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_USER: ${DB_USERNAME}
      POSTGRES_DB: ${DB_DATABASE_NAME}
      POSTGRES_INITDB_ARGS: '--data-checksums'
    volumes:
      # Do not edit the next line. If you want to change the database storage location on your system, edit the value of DB_DATA_LOCATION in the .env file
      - ${DB_DATA_LOCATION}:/var/lib/postgresql/data
    healthcheck:
      test: >-
        pg_isready --dbname="$${POSTGRES_DB}" --username="$${POSTGRES_USER}" || exit 1; Chksum="$$(psql --dbname="$${POSTGRES_DB}" --username="$${POSTGRES_USER}
" --tuples-only --no-align --command='SELECT COALESCE(SUM(checksum_failures), 0) FROM pg_stat_database')"; echo "checksum failure count is $$Chksum"; [ "$$Chksu
m" = '0' ] || exit 1
      interval: 5m
      start_interval: 30s
      start_period: 5m
    command: >-
      postgres -c shared_preload_libraries=vectors.so -c 'search_path="$$user", public, vectors' -c logging_collector=on -c max_wal_size=2GB -c shared_buffers=5
12MB -c wal_compression=on
    restart: always

volumes:
  model-cache:

Your .env content

# You can find documentation for all the supported env variables at https://immich.app/docs/install/environment-variables

# The location where your uploaded files are stored
UPLOAD_LOCATION=./library

# The location where your database files are stored. Network shares are not supported for the database
DB_DATA_LOCATION=./postgres

# To set a timezone, uncomment the next line and change Etc/UTC to a TZ identifier from this list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones#
List
# TZ=Etc/UTC

# The Immich version to use. You can pin this to a specific version like "v1.71.0"
IMMICH_VERSION=release

# Connection secret for postgres. You should change it to a random password
# Please use only the characters `A-Za-z0-9`, without special characters or spaces
DB_PASSWORD=XXXXXXX # THIS IS SET, JUST NOT SHOWING HERE

# The values below this line do not need to be changed
###################################################################################
DB_USERNAME=postgres
DB_DATABASE_NAME=immich

Reproduction steps

  1. Administration
  2. Jobs
  3. Sidecar Metadata -> Refresh

...

It starts, processes maybe about half, then hangs. I've had this happen on other jobs as well.

Relevant log output


Additional information

No response

Originally created by @davejhahn on GitHub (May 17, 2025). ### I have searched the existing issues, both open and closed, to make sure this is not a duplicate report. - [x] Yes ### The bug I have a large collection, approximately 125K images, so maybe this is something that surfaces in this scenario, but I frequently have jobs that just get stuck. e.g. Extract Metadata, it will process for a while, I see the counts going down, and then it just hangs, as do any other jobs. The only way I can get past it is to stop the service (docker compose down) and restart. Sometimes it works then, other times it repeats. I've streamed the log and do not see anything being logged that indicates what might be going on. Nothing appears in logs, at least the way I am looking at them: `docker logs --follow immich_server` ### The OS that Immich Server is running on Ubuntu 22.04 ### Version of Immich Server v1.32.3 ### Version of Immich Mobile App 1.132.3.build.205 ### Platform with the issue - [ ] Server - [x] Web - [ ] Mobile ### Your docker-compose.yml content ```YAML # # WARNING: To install Immich, follow our guide: https://immich.app/docs/install/docker-compose # # Make sure to use the docker-compose.yml of the current release: # # https://github.com/immich-app/immich/releases/latest/download/docker-compose.yml # # The compose file on main may not be compatible with the latest release. name: immich services: immich-server: container_name: immich_server image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release} # extends: # file: hwaccel.transcoding.yml # service: cpu # set to one of [nvenc, quicksync, rkmpp, vaapi, vaapi-wsl] for accelerated transcoding volumes: # Do not edit the next line. If you want to change the media storage location on your system, edit the value of UPLOAD_LOCATION in the .env file - ${UPLOAD_LOCATION}:/usr/src/app/upload - /etc/localtime:/etc/localtime:ro - /images:/images env_file: - .env ports: - '2283:2283' depends_on: - redis - database restart: always healthcheck: disable: false immich-machine-learning: container_name: immich_machine_learning # For hardware acceleration, add one of -[armnn, cuda, rocm, openvino, rknn] to the image tag. # Example tag: ${IMMICH_VERSION:-release}-cuda image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release} # extends: # uncomment this section for hardware acceleration - see https://immich.app/docs/features/ml-hardware-acceleration # file: hwaccel.ml.yml # service: cpu # set to one of [armnn, cuda, rocm, openvino, openvino-wsl, rknn] for accelerated inference - use the `-wsl` version for WSL2 where applica ble volumes: - model-cache:/cache env_file: - .env restart: always healthcheck: disable: false redis: container_name: immich_redis image: docker.io/valkey/valkey:8-bookworm@sha256:42cba146593a5ea9a622002c1b7cba5da7be248650cbb64ecb9c6c33d29794b1 healthcheck: test: redis-cli ping || exit 1 restart: always database: container_name: immich_postgres image: docker.io/tensorchord/pgvecto-rs:pg14-v0.2.0@sha256:739cdd626151ff1f796dc95a6591b55a714f341c737e27f045019ceabf8e8c52 environment: POSTGRES_PASSWORD: ${DB_PASSWORD} POSTGRES_USER: ${DB_USERNAME} POSTGRES_DB: ${DB_DATABASE_NAME} POSTGRES_INITDB_ARGS: '--data-checksums' volumes: # Do not edit the next line. If you want to change the database storage location on your system, edit the value of DB_DATA_LOCATION in the .env file - ${DB_DATA_LOCATION}:/var/lib/postgresql/data healthcheck: test: >- pg_isready --dbname="$${POSTGRES_DB}" --username="$${POSTGRES_USER}" || exit 1; Chksum="$$(psql --dbname="$${POSTGRES_DB}" --username="$${POSTGRES_USER} " --tuples-only --no-align --command='SELECT COALESCE(SUM(checksum_failures), 0) FROM pg_stat_database')"; echo "checksum failure count is $$Chksum"; [ "$$Chksu m" = '0' ] || exit 1 interval: 5m start_interval: 30s start_period: 5m command: >- postgres -c shared_preload_libraries=vectors.so -c 'search_path="$$user", public, vectors' -c logging_collector=on -c max_wal_size=2GB -c shared_buffers=5 12MB -c wal_compression=on restart: always volumes: model-cache: ``` ### Your .env content ```Shell # You can find documentation for all the supported env variables at https://immich.app/docs/install/environment-variables # The location where your uploaded files are stored UPLOAD_LOCATION=./library # The location where your database files are stored. Network shares are not supported for the database DB_DATA_LOCATION=./postgres # To set a timezone, uncomment the next line and change Etc/UTC to a TZ identifier from this list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones# List # TZ=Etc/UTC # The Immich version to use. You can pin this to a specific version like "v1.71.0" IMMICH_VERSION=release # Connection secret for postgres. You should change it to a random password # Please use only the characters `A-Za-z0-9`, without special characters or spaces DB_PASSWORD=XXXXXXX # THIS IS SET, JUST NOT SHOWING HERE # The values below this line do not need to be changed ################################################################################### DB_USERNAME=postgres DB_DATABASE_NAME=immich ``` ### Reproduction steps 1. Administration 2. Jobs 3. Sidecar Metadata -> Refresh ... It starts, processes maybe about half, then hangs. I've had this happen on other jobs as well. ### Relevant log output ```shell ``` ### Additional information _No response_
OVERLORD added the 🗄️server label 2026-02-05 11:56:44 +03:00
Author
Owner

@bo0tzz commented on GitHub (May 17, 2025):

What sort of system are you running on, which jobs are hanging (and what does this hanging look like), how many assets do you have, etc? Consider that some jobs will queue new jobs after they finish, so numbers not going down doesn't necessarily mean things are hung.

@bo0tzz commented on GitHub (May 17, 2025): What sort of system are you running on, which jobs are hanging (and what does this hanging look like), how many assets do you have, etc? Consider that some jobs will queue new jobs after they finish, so numbers not going down doesn't necessarily mean things are hung.
Author
Owner

@davejhahn commented on GitHub (May 17, 2025):

I am running on Linux, I have have a RTX 4060 Ti w/16GB, and a i7-14700F with 82GB of memory.

I just tried something and it seemed to stop happening. I had updated the concurrency setting at some point. I reduced it to 5, and it seems that it repeatedly works. Not sure how the concurrency actually works but my original thought was I have a relatively powerful system, increase it, apparently that was not a good idea?

@davejhahn commented on GitHub (May 17, 2025): I am running on Linux, I have have a RTX 4060 Ti w/16GB, and a i7-14700F with 82GB of memory. I just tried something and it seemed to stop happening. I had updated the concurrency setting at some point. I reduced it to 5, and it seems that it repeatedly works. Not sure how the concurrency actually works but my original thought was I have a relatively powerful system, increase it, apparently that was not a good idea?
Author
Owner

@davejhahn commented on GitHub (May 17, 2025):

Related to the not processing, it is definitely not when this happens, I've left it in this state overnight, and nothing changes.

@davejhahn commented on GitHub (May 17, 2025): Related to the not processing, it is definitely not when this happens, I've left it in this state overnight, and nothing changes.
Author
Owner

@bo0tzz commented on GitHub (May 17, 2025):

I had updated the concurrency setting at some point. I reduced it to 5, and it seems that it repeatedly works. Not sure how the concurrency actually works but my original thought was I have a relatively powerful system, increase it, apparently that was not a good idea?

Idk what concurrency you were originally running, but we often see people set it way too high which definitely causes problems. Even on a fairly powerful system the numbers don't go as high as you'd expect, especially on slower storage.

@bo0tzz commented on GitHub (May 17, 2025): > I had updated the concurrency setting at some point. I reduced it to 5, and it seems that it repeatedly works. Not sure how the concurrency actually works but my original thought was I have a relatively powerful system, increase it, apparently that was not a good idea? Idk what concurrency you were originally running, but we often see people set it way too high which definitely causes problems. Even on a fairly powerful system the numbers don't go as high as you'd expect, especially on slower storage.
Author
Owner

@bigZos commented on GitHub (May 19, 2025):

something similar happened to me. running via docker desktop on Windows (WSL2 backend) with a library of ~90k photos/videos and default concurrency settings for its jobs.

I've experienced situations where scheduled Immich machine learning tasks lead to very high system memory usage. there was an instance where docker desktop's backend process itself retained over 50GB of RAM even after attempting to stop the immich containers. This memory was only released after a full restart of docker desktop.

this sustained memory pressure eventually led to system-wide instability and a crash about 12hrs later (i have it set to scan external libraries every 6 hours) (in my case, an NVIDIA driver failed, but this was after hours of low memory warnings and other services failing due to resource exhaustion).

while my jobs didn't "hang" in the same way described, the failure of Docker to reclaim substantial resources after immich tasks could be a contributing to job-related issues or resource exhaustion on windows. Hope this helps

@bigZos commented on GitHub (May 19, 2025): something similar happened to me. running via docker desktop on Windows (WSL2 backend) with a library of ~90k photos/videos and default concurrency settings for its jobs. I've experienced situations where scheduled Immich machine learning tasks lead to very high system memory usage. there was an instance where docker desktop's backend process itself retained over 50GB of RAM even after attempting to stop the immich containers. This memory was only released after a full restart of docker desktop. this sustained memory pressure eventually led to system-wide instability and a crash about 12hrs later (i have it set to scan external libraries every 6 hours) (in my case, an NVIDIA driver failed, but this was after hours of low memory warnings and other services failing due to resource exhaustion). while my jobs didn't "hang" in the same way described, the failure of Docker to reclaim substantial resources after immich tasks could be a contributing to job-related issues or resource exhaustion on windows. Hope this helps
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: immich-app/immich#6019