Files
microfish/Dockerfile
Kunthawat Greethong 851ed65f45 build: resolve torch as CPU-only to eliminate CUDA download
camel-oasis (transitively via sentence-transformers) hard-imports torch in
oasis/social_platform/recsys.py, so torch cannot be removed while OASIS
simulation is a feature. On linux-x86_64 the default PyPI torch wheel is the
CUDA build, which dragged ~several GB of nvidia-* packages into the image
even though this server has no GPU and all LLM + embedding calls go through
an API (generate_post_vector_openai).

Fix: force torch to resolve from PyTorch's CPU-only index via [tool.uv]:
- override-dependencies torch==2.13.0, [[tool.uv.index]] pytorch-cpu, and
  [tool.uv.sources] torch={index=pytorch-cpu}.

Result after re-lock: all nvidia-* + triton packages removed (0 remaining in
uv.lock), torch 2.9.1 -> 2.13.0+cpu. Verified: torch/sentence_transformers/
oasis/from app import create_app all import fine with CPU torch
(cuda: False); backend suite 201 passed. Dockerfile keeps an import-time
verify after uv sync instead of the now-unneeded nvidia uninstall step.
2026-08-31 21:11:32 +07:00

75 lines
4.3 KiB
Docker

# ============================================================
# CrowdSight production image (multi-service, EasyPanel-buildable)
#
# Services inside one container (supervisord):
# - web: nginx serving the built SPA, proxying /api -> gunicorn :5001
# - backend: gunicorn WSGI (wsgi:app) on 0.0.0.0:5001
# - worker: durable PollingWorker (backend/worker.py)
# ============================================================
# ---- Stage 1: build the frontend SPA ----
FROM node:20 AS frontend-build
WORKDIR /build
COPY package.json package-lock.json* ./
COPY frontend/package.json frontend/package-lock.json* ./frontend/
# install root + frontend deps
RUN npm ci 2>/dev/null || true; npm ci --prefix frontend || true
COPY locales ./locales
COPY frontend ./frontend
RUN npm run build --prefix frontend
# ---- Stage 2: runtime (python + nginx + supervisord) ----
FROM python:3.11-slim AS runtime
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
PYTHONPATH=/app/backend
RUN apt-get update \
&& apt-get install -y --no-install-recommends nginx supervisor \
&& rm -rf /var/lib/apt/lists/*
# uv runtime
COPY --from=ghcr.io/astral-sh/uv:0.9.26 /uv /uvx /bin/
WORKDIR /app
# Copy project source FIRST so the installed deps match the shipped
# pyproject.toml/uv.lock exactly (avoids `uv run` re-syncing at container
# runtime, which previously rebuilt the project every time the entrypoint
# ran and missed the psycopg Postgres driver -> "NoSuchModuleError:
# sqlalchemy.dialects:postgres").
COPY backend ./backend
COPY locales ./locales
COPY package.json ./
# Install backend deps against the final pyproject.toml/uv.lock. torch resolves
# to a CPU-only wheel (via [tool.uv] -> pytorch-cpu index in pyproject.toml), so
# no nvidia-* CUDA packages are downloaded. Verify the env still imports the
# Postgres driver, torch, sentence-transformers (used by camel-oasis recsys) and
# the app, so the build fails loudly rather than at container runtime.
RUN cd backend && uv sync --frozen --no-dev \
&& uv run --frozen python -c "import psycopg, torch; import sentence_transformers; from app import create_app; print('deps-after-sync-ok')" \
|| (echo "FATAL: dependency import failed (psycopg/torch/sentence-transformers/app). Revisit dependency resolution." >&2 && exit 1)
# Make the migration-runner entrypoint executable
RUN chmod +x /app/backend/docker_entrypoint.sh
# Copy built SPA into nginx web root
COPY --from=frontend-build /build/frontend/dist /usr/share/nginx/html
# nginx config: SPA + /api proxy to gunicorn
RUN echo 'server {\n listen 8080;\n server_name _;\n root /usr/share/nginx/html;\n index index.html;\n location / { try_files $uri $uri/ /index.html; }\n location /api/ {\n proxy_pass http://127.0.0.1:5001;\n proxy_set_header Host $host;\n proxy_set_header X-Real-IP $remote_addr;\n proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;\n proxy_set_header X-Forwarded-Proto $scheme;\n }\n}\n' > /etc/nginx/sites-available/crowdsight \
&& ln -sf /etc/nginx/sites-available/crowdsight /etc/nginx/sites-enabled/crowdsight \
&& rm -f /etc/nginx/sites-enabled/default
# supervisor: run nginx + gunicorn + worker (migrations already run by entrypoint)
RUN echo '[supervisord]\nnodaemon=true\nlogfile=/var/log/supervisor/supervisord.log\npidfile=/var/run/supervisord.pid\n\n[program:nginx]\ncommand=/usr/sbin/nginx -g "daemon off;"\nautostart=true\nautorestart=true\nstdout_logfile=/dev/stdout\nstdout_logfile_maxbytes=0\nstderr_logfile=/dev/stderr\nstderr_logfile_maxbytes=0\n\n[program:backend]\ncommand=/bin/bash -c "cd /app/backend && uv run --frozen gunicorn -w 2 -b 0.0.0.0:5001 --timeout 120 wsgi:app"\nautostart=true\nautorestart=true\nstdout_logfile=/dev/stdout\nstdout_logfile_maxbytes=0\nstderr_logfile=/dev/stderr\nstderr_logfile_maxbytes=0\n\n[program:worker]\ncommand=/bin/bash -c "cd /app/backend && uv run --frozen python worker.py --poll-interval 5"\nautostart=true\nautorestart=true\nstartsecs=2\nstartretries=5\nstdout_logfile=/dev/stdout\nstdout_logfile_maxbytes=0\nstderr_logfile=/dev/stderr\nstderr_logfile_maxbytes=0\n' > /etc/supervisor/conf.d/crowdsight.conf
EXPOSE 8080 5001
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s \
CMD python -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:5001/health', timeout=4)" || exit 1
CMD ["/app/backend/docker_entrypoint.sh"]