Elevate MiroFish/CrowdSight from single-container dev to a SaaS foundation: - Local memory backend (Zep-compatible): memory services/models, local graph builder + updater, AgentActivity seam, import-boundary isolation; Zep stays default, local is opt-in behind MEMORY_BACKEND. Semantic parity not yet proven. - Durable product persistence: projects/simulations/reports schema (migration 0007) + tenant/owner-scoped ProductRepository + dual-write + scoped_project read-first + ArtifactStore abstraction; durable JobQueue + worker.py. - SaaS hardening: durable RateLimiter (wired to login), UsageService (LLM accounting), redacted AuditService, idempotency, CORS allowlist, safe API errors, single-use PasswordResetService + endpoints (covers invite-pending). - Exactly 3 roles (super_admin/admin/user) with tenant authz policy. - Admin UI: GET/POST/PATCH /api/admin/users + GET/PUT /api/admin/settings (super-admin only, encrypted/masked); AdminView.vue + SettingsView.vue with admin/super-admin route guards, th/en i18n. - Production deploy topology: multi-stage Dockerfile (frontend build + gunicorn wsgi + nginx SPA-proxy + supervisord worker), backend/wsgi.py, gunicorn dep. Backend 197 passed; frontend 10 tests + build green. ruff unavailable (gap). No commit of credentials; secrets handled via env/.env.example. Deferred: Zep semantic A/B parity, object storage cutover, mobile QA, EasyPanel container build of deploy topology.
70 lines
2.2 KiB
Python
70 lines
2.2 KiB
Python
"""Durable LLM usage/cost accounting service.
|
|
|
|
Records per-organization, per-user LLM usage without storing any prompt content
|
|
or secrets. A simple default cost estimate (input/output per-token) is applied
|
|
and can be overridden by a rate table later.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Optional
|
|
|
|
from sqlalchemy.orm import Session
|
|
|
|
from ..models.usage import UsageEvent
|
|
|
|
# Default per-1K token cost estimates (USD); a rate table can supersede later.
|
|
_DEFAULT_INPUT_RATE_PER_1K = 0.0025
|
|
_DEFAULT_OUTPUT_RATE_PER_1K = 0.0100
|
|
|
|
|
|
class UsageService:
|
|
def __init__(self, session: Session):
|
|
self.session = session
|
|
|
|
def record_event(
|
|
self,
|
|
*,
|
|
organization_id: str,
|
|
user_id: Optional[str],
|
|
operation: str,
|
|
model: Optional[str] = None,
|
|
input_tokens: int = 0,
|
|
output_tokens: int = 0,
|
|
) -> str:
|
|
if not isinstance(organization_id, str) or not organization_id:
|
|
raise ValueError("organization_id_required")
|
|
cost = (
|
|
(input_tokens / 1000) * _DEFAULT_INPUT_RATE_PER_1K
|
|
+ (output_tokens / 1000) * _DEFAULT_OUTPUT_RATE_PER_1K
|
|
)
|
|
event = UsageEvent(
|
|
organization_id=organization_id,
|
|
user_id=user_id,
|
|
operation=operation,
|
|
model=model,
|
|
input_tokens=int(input_tokens or 0),
|
|
output_tokens=int(output_tokens or 0),
|
|
estimated_cost=round(cost, 6),
|
|
)
|
|
self.session.add(event)
|
|
self.session.flush()
|
|
return event.id
|
|
|
|
def list_events(self, *, organization_id: str, limit: int = 100) -> list[UsageEvent]:
|
|
return (
|
|
self.session.query(UsageEvent)
|
|
.filter(UsageEvent.organization_id == organization_id)
|
|
.order_by(UsageEvent.created_at.desc())
|
|
.limit(min(max(int(limit), 1), 1000))
|
|
.all()
|
|
)
|
|
|
|
def total_cost(self, *, organization_id: str) -> float:
|
|
rows = (
|
|
self.session.query(UsageEvent)
|
|
.filter(UsageEvent.organization_id == organization_id)
|
|
.all()
|
|
)
|
|
return round(sum(row.estimated_cost for row in rows), 6)
|