I put LLM and AI-agent systems into production on serverless and rented GPUs. Sub-200ms p95 on 7B to 70B models, served with vLLM and GPTQ/AWQ quantization, at roughly 60% of what the managed APIs cost.
Freelance, based in Bangladesh · Email · GitHub · LinkedIn · CV (PDF)
Upwork Top Rated · 100% Job Success Score
I trained as an electrical engineer, so I learned computers from the bottom up: semiconductor junctions, then transistors, then the machines built out of them. I now work at the other end of that stack, keeping GPU clusters fed. Most people debugging a CUDA out-of-memory error are reasoning about an abstraction. I'm reasoning about a device with a memory hierarchy, a thermal budget, and a bus that has opinions.
ODS assumed every service spoke HTTP, so anything exposing only TCP or a CLI was marked dead and vanished from the dashboard without an error anyone would see. I threaded a health_type field through the schema, catalog generator, dashboard API and shell scripts, with defaults that inferred correctly so dozens of existing manifests kept working untouched.
TCP-only and CLI-only services stay visible instead of being reported dead. 902 lines added across the stack, with no change required to any existing manifest. Still under review.
Multi-agent systems waste money re-serializing thoughts into English so the next agent can parse them back. I extended the LatentMAS paper (ICML 2026 Spotlight) with role-specialized LoRA adapters that hot-swap at runtime, so one base model plays four agents without four sets of weights in VRAM. The authors added it to their README as a community extension. The repo is equally clear about what it isn't: PEFT adapter management, not true S-LoRA.
12% higher accuracy, 2.7x faster inference, 63.6% fewer tokens and 75% less VRAM, from four-role adapter swapping on a single Qwen 2.5 base.
GPU marketplaces hand you a box that is technically the card you paid for and broken in a dozen quiet ways. I wrote a hardened multi-phase installer for it, with GPU-tier detection and 28 documented host-environment failure modes and their fixes. The documentation was most of the value. Anyone can write the happy path.
A 6,250-line P2P GPU deployment toolkit for Vast.ai, covering driver-mismatch repair and 28 host-failure modes. Still under review.
Wan 2.2 text-to-video, image-to-video and speech-to-video, run as distributed inference with FSDP on serverless clusters. Topology-aware placement decides which cards talk to each other, and per-GPU VRAM capping keeps a single greedy rank from taking the whole job down.
3x throughput against the single-GPU baseline.
End-to-end serverless workers for FLUX.2, Whisper and custom diffusion models. Dockerized, CI/CD, autoscaled down when no one is calling them. A LoRA training pipeline sits alongside, so a client can fine-tune and serve from the same setup.
$0 idle cost between requests.
More, including an MCP server that brokers GPU compute for agents, at github.com/Arifuzzamanjoy.
aimclub/OSA. OSA writes CONTRIBUTING, SECURITY and CODE_OF_CONDUCT for a repository but nothing a machine reads, so I added a codemeta.json generator. That is the JSON-LD descriptor Zenodo, Software Heritage and most FAIR-software checklists index against. It falls back to pyproject.toml for what the Git host doesn't expose, handles PEP 621 and Poetry author formats, and maps license names to SPDX URLs.
vLLM omni-modality. build_engine_args_dict was mutating the caller's stage_config.engine_args, so engine args leaked from one pipeline stage into the next. A one-line deepcopy fix and 145 lines of tests covering idempotency and nested-dict isolation. The fix is trivial. Proving it stays fixed isn't.
Last updated 4 September 2026.
Send the model, your latency target, and the budget. You'll get a straight answer about what it takes, including when the answer is that you don't need me. I do my best work when the thing being built has some reason to exist beyond a funding round.
Hire me on Upwork joy.apee@gmail.com
Academic reader? My publications, research interests and graduate-study plans are on the research page.