Coding agents need a delivery system
A field report on coordinating Claude Code, Codex, review, repository state, and human authority while shipping a real open-source project.
I help teams turn messy data and half-working AI prototypes into reliable products that people can use.
I map the workflow before choosing models, vendors, or architecture.
I get useful pilots in front of users, then measure and improve them.
Teams should know what the system does, where it fails, and which trade-offs they accept.
Most AI work fails in the handoff between idea, data, product, and production. I work across that handoff.
I design and ship LLM applications: RAG chatbots, knowledge assistants, and LLM eval suites tailored to your data and use case. Quick wins that move the needle, from early pilots to 10k-user rollouts.
I turn messy data into structured, usable datasets and build semantic search and recommender systems on top, using LLMs for feature engineering, enrichment, and large-scale dataset creation.
I help teams pick the right AI bets, build a realistic roadmap, align stakeholders, and grow internal capability while staying close enough to the implementation to keep the advice honest.
I have worked across affiliate technology, retail, SaaS, and enterprise AI, usually in the messy middle between data science, product, and engineering.
Architecting and delivering production AI systems across RAG, agents, evaluation, and reusable platform capabilities.
Built an affiliate-marketing assistant delivered through Slack and Microsoft Teams.
Scoped and delivered data science projects for retail, then explored early generative AI use cases.
Built customer analytics and NLP systems for relationship marketing and banking teams.
I use the blog to think in public about LLM engineering, AI platforms, evaluation, and probabilistic software.
A field report on coordinating Claude Code, Codex, review, repository state, and human authority while shipping a real open-source project.
Published · Updated
Start from delivery bottlenecks, then build the context, evaluation, operations, and reusable primitives that help teams ship.
Classic software is deterministic; LLM applications add a probabilistic core that needs habits from data science and software engineering.
Practical walkthroughs and references that show how the systems behave.
“Othman is a first-class Data Scientist and engineering team member.
He has an excellent work ethic, and unwavering dedication to his colleagues and work.
I had the pleasure of working with him on a new generative AI solution designed to meet the needs of thousands of customers.
At the time of writing (March 2025), building generative AI applications can be highly complex, with a multitude of unknowns and unpredictable outcomes due to the probabilistic nature of generative AI models.
As a key member in our team, Othman designed and implemented our evaluations methodology and technology stack. His work has been invaluable to understanding the performance and optimization requirements for our solution. Beyond his technical expertise, Othman is a superb communicator, adept at speaking to counterparts at all levels in our organization, from the most technically proficient, to critical non-technical stakeholders.
I am well aware that Othman's skillsets extend far beyond the remit that we have worked together. Beyond the evaluations work, he also conducts deep quantitative analysis of unstructured data to obtain meaningful business insights.
I highly recommend Othman as an experienced and trusted colleague.”
“Othman is in the top 1% of engineers and humans I have ever worked with. If you are lucky enough to have the opportunity to do so, work with Othman.”
I’m a freelance AI and data science consultant with eight years of experience across affiliate technology, retail, SaaS, and enterprise AI. I design and ship LLM products, data pipelines, and recommender systems end to end, from messy data to production.
I sit between strategy and engineering, helping teams choose the right AI bets and actually get them shipped.
Off the keyboard, I’m a distance runner and philosophy-language nerd who loves quirky questions.
I like working with teams who want straight thinking, practical systems, and a light enough tone that the work stays fun.
Bring the messy version. We can talk through the workflow, the data, the risks, and what a useful first step would look like.
Book a 15-min call