I build AI systems that reach production — and stay there.
A decade of production engineering — Salesforce, dLocal, The Sandbox. Today, AI that handles 15,000+ conversations a month at 90%+ accuracy.
Currently AI Engineer at SignalCore AI, building autonomous decision systems for enterprise.
A backend engineer who moved to the AI layer — and ships it.
I’m Ramiro González, a software engineer with a decade building backend systems for companies like Salesforce, dLocal and The Sandbox. For the last few years I’ve worked entirely on AI — taking language models out of the demo and into production, where real people and real load are what break them.
I don’t sell packaged services. I take on a project and go all in: I map how the work happens today, get something running in weeks, and hand you code you own outright. Based in Florianópolis, Brazil, I work across US and European hours.
Five capabilities — each with something shipped behind it.
I’m one engineer, not an agency — so I go deep, not wide. This is the ground I cover when I take a project on.
RAG architecture
90%+ answer relevance with citations; 65% fewer hallucinations via semantic chunking and hybrid search.
Agent orchestration
10+ agents in production, coordinating tools, memory and human handoff.
LLM systems at volume
15,000+ conversations a month, holding accuracy under real production load.
Evaluation & guardrails
Every system ships with LLM-as-judge evals, cost and latency monitoring, and built-in escalation.
Backend at scale
120K monthly active users at The Sandbox; cross-border payments infrastructure at dLocal.
What working together actually looks like.
We start with a 15-minute scoping call. Free. If it isn’t a fit, I’ll tell you on that call.
I map the process before I write code. Most AI projects fail by automating a broken process faster. First I want to see how the work happens today.
Something runs by week three or four, even if it’s narrow. A system you can use badly beats one you can only discuss.
I’ll tell you what the AI should not do. Every project has a line where the machine stops and a person starts. Drawing it early is most of the difference between a production system and a demo.
You own the code. Repository, infrastructure, prompts — all of it, from day one.
Based in Florianópolis (UTC−3) · hours overlap the US & most of Europe · English · Portuguese · Spanish
Niches I’ve already shipped into.
If your world is on this list, I’ve built for it before — as a founder’s first AI hire, or as the engineer behind the product.
An always-on AI SDR for aesthetic clinics.
Turns WhatsApp leads into booked consultations — and hands closers only the ones ready to buy.
The agent’s job is to bring twenty-five qualified people to a human closer. Not to close them.
The clinic — I’ll call them Concierge Doctor here — ran entirely on WhatsApp. When ten leads arrived at once, the tenth waited two hours. Aesthetic leads are impulse-driven — a two-hour reply is a dead lead.
The clinic answered about a quarter of the inbound it could generate, so it couldn’t buy traffic at all — more ads only made the queue longer. Bookings sat at five a month.
A WhatsApp concierge that qualifies, nurtures and follows up 24/7. Behind it, an operations cockpit that scores every lead, writes a brief for the human closer, and books the consultation.
- —500 leads handled
- —First response under 30 seconds, regardless of volume
- —Booked consultations 5 → 25 per month
At the clinic’s R$5,000 average ticket, twenty extra consultations is roughly R$100,000/month in new consultation pipeline.
The agent hands to a human on exactly three triggers: it fails to resolve an objection after three attempts, the lead asks for a person, or it detects it lacks the information to answer correctly.
The third one matters most. Most clinic bots fail by improvising through a knowledge gap — an invented price, a promised result. That costs trust, and it can cost a treatment. So the system recognises the edge of its own knowledge and stops there.
The same system, running on a made-up lead.
WhatsApp is just the transport. The qualification, scoring and closer-brief logic is the product — it runs the same on a chat, a form or a call transcript.
Screens from the system in production — the real interface behind the demo above.
Support that answers from the docs — and knows when it can’t.
A custom retrieval agent embedded in an ed-tech product. Instant, cited answers — and a clean escalation the moment it isn’t sure.
* Illustrative figures — final production numbers pending.
The agent answers what the docs already know — and routes everything else to a person, fast.
A growing ed-tech company — small team, thousands of teachers and students. Support was a few people answering the same questions across email and chat. The help centre was scattered, so the answer you got depended on who replied.
As sign-ups grew, response times slipped. Hiring more agents scales cost, not consistency — and the fastest-growing queue was the same fifty questions asked a hundred ways.
A custom retrieval-augmented agent, embedded in the product as a widget. It answers only from a curated, standardised knowledge base, always with citations, and escalates the moment retrieval confidence drops or the question is out of scope.
- —71% of conversations resolved end-to-end without a human
- —90%+ answer accuracy against a golden Q&A set
- —First response 4h → <5s
- —Support scaled with sign-ups, no new headcount
Every answer is grounded in an approved source and shows its citations. When the top retrieval score falls below threshold — or the question is out of scope, like pricing or anything account-specific — the agent doesn’t improvise. It escalates with the full conversation attached.
A confidently wrong support answer is worse than a slow one: it teaches the user something false and quietly erodes trust in the product.
One question, traced end to end.
The widget is what the user sees. The trace beside it is the retrieval pipeline that produced the answer — and the guardrail that escalates instead of guessing.
Decision memory for architecture firms.
It turns the meetings that already happen into a searchable record of every decision — who made it, why, and what it changed.
Pre-launch · in pilot with souBIM, an architecture firm in Florianópolis
souBIM runs three or more projects at once. Every material change, structural revision and budget call happened in a Google Meet — and lived only in the head of Gabriela, the director. Nothing was written down.
Then Gabriela scheduled two to three months of maternity leave. Around fifty decisions per project would be made without her. It wasn’t a leave problem — it was a scaling problem: the firm couldn’t grow past one person’s memory.
Every decision, filterable and attributed — surfaced as a project timeline and a “while you were gone” digest.
- —150–200 decisions a month captured with zero meeting friction
- —Meeting ends → decision logged in under 4 hours
The system extracts and structures decisions — it never makes or approves them. It records consensus and dissent instead of smoothing them into a clean answer, and it scores its own confidence on every extraction.
When a decision is low-consensus or contradicts an earlier one, it flags it for the director rather than asserting a tidy record. A confidently wrong decision log is worse than a flagged uncertain one — in a building, trust is load-bearing.
The system’s job is to give a director back her memory. Not to make her decisions.
Other projects, across other worlds.
A few more at different stages — two shipping right now, and one that won a competition and then shipped to paying clients.
AI Librarian for a financial-education platform
An in-product agent that answers from the platform’s videos and lessons, guiding members through their learning journey toward real results.
Project calls, turned into a searchable knowledge base
Ingested an architecture firm’s call recordings and turned them into a searchable content platform. Delivered on Upwork with a five-star client review.
“Deeply impressed with his knowledge of RAG systems, clear communication, and depth of architectural understanding.”
Digital clones of how a person decides and speaks
Chain-of-thought and few-shot systems with ElevenLabs voice, replicating a person’s decisions and style. First of 200 at Academia Lendária, then shipped to three paying clients.
A decade of production engineering.
Independent, contract engagements — from cloud CRM and cross-border payments to the AI systems above.