Full Stack AI Developer (Contract)
<p><strong>Full-time freelance · 2 months · Extendable month to month<span class="ql-cursor"></span></strong></p><p>We build full-stack applications where the core logic is an LLM workflow or an agent. You will build them end to end: the interface people use, the backend that orchestrates the models, and the engineering that makes the whole thing reliable enough to put in front of a client.</p><p><br></p><p><br></p>What you'll own<ul><li><strong>Complete application builds.</strong> From first prototype to deployable product, across frontend, API, orchestration, data and deployment.</li><li><strong>Agentic backends in LangGraph.</strong> Graph design, state management, tool calling, branching, retries, checkpointing, and human-in-the-loop interrupts where a person needs to approve what the agent is about to do.</li><li><strong>Frontends designed for AI.</strong> Streaming responses, visible intermediate steps, graceful handling of slow or partial failures, and controls that let users correct or approve model output. A twenty-second spinner is a broken product.</li><li><strong>Making model behaviour dependable.</strong> Validated structured outputs, fallbacks, guardrails, and evals that tell you whether a change actually made things better.</li><li><strong>Connecting to the real world.</strong> Third-party APIs, databases, documents and internal systems, with authentication and permissions handled properly.</li><li><strong>Fast, visible iteration.</strong> Short cycles with working demos, and the judgement to suggest a smaller scope when the original one won't ship in time.</li></ul><p><br></p><p><br></p>What we're looking for<ul><li><strong>Full-stack apps you have shipped with an LLM at the core.</strong> Send us links or repositories. We want to see something that does more than wrap a chat box around an API call.</li><li><strong>Hands-on LangGraph experience, in Python or JavaScript.</strong> You can explain how state moves through your graph, what happens when a node fails halfway through, and how you resume a run.</li><li><strong>Solid frontend engineering.</strong> React (including Next.js) with TypeScript, or Flutter, including streaming interfaces and real-time updates.</li><li><strong>Solid backend engineering.</strong> Python or TypeScript/Node, a relational database such as Postgres, authentication, and asynchronous processing.</li><li><strong>A working command of LLM fundamentals.</strong> How context windows and token costs shape design; structured output and tool calling; RAG and when it isn't the answer; choosing between reasoning and instruct models, or large and small ones, on cost and latency; why prompt injection is solved with permissions rather than prompts; and how to evaluate output that is never quite the same twice.</li><li><strong>Comfort with ambiguity and pace.</strong> Requirements will move. You should be able to keep shipping while they do.</li></ul><p><br></p><p><br></p>Signals that would move you to the top of the list<ul><li><strong>MCP experience,</strong> whether building servers or wiring agents to consume them.</li><li><strong>Observability and evals in practice.</strong> LangSmith, Langfuse or your own tracing, and a test set you actually run.</li><li><strong>Multi-provider work.</strong> OpenAI, Anthropic, Gemini or open-weight models, with routing and fallback between them.</li><li><strong>Retrieval done well.</strong> Vector stores such as pgvector or Qdrant, hybrid search, reranking, and retrieval quality you have measured.</li><li><strong>Deployment ownership.</strong> Docker, a cloud platform and CI/CD, so a build can go live without a separate handoff.</li><li><strong>Product sense.</strong> You are comfortable talking to clients and users, and you push back when a feature is expensive to build and unlikely to matter.</li></ul><p><br></p><p><br></p>What this role is not<p>It is not research, and you will not be training models. It is not prompt engineering in a notebook. It is not frontend-only or backend-only work either. You will own both halves of applications that real people use, with the model treated as one unreliable component among many.</p>