The quality of a product is rarely defined by the features you build. It's defined by the decisions you make long before the first line of code is written.
I'm Gabo. I work at the intersection of strategy, technology, and business, turning ambiguity into a direction a team can actually build toward. Lately that's meant deciding which AI platform gets which job, and why "pick one winner" is usually the wrong question.
I think a good product conversation should feel like a good brainstorm: structured enough to go somewhere, loose enough that a good idea can still surprise you. I'd rather ask an uncomfortable question early than write a polished roadmap around the wrong assumption.
Most of what follows isn't about features I shipped. It's about the calls I made before shipping was even on the table, and what I'd do differently now that I've seen how they played out.
Product management for real clients at Imaginex Studio, from a 20-person healthcare build to three-week sprints. Each one opens with the decision, not the feature list.
A 90-day pilot to decide whether AI agents belonged inside a global fashion brand, and if so, on which platform. Five vendors, ~30 real use cases across six departments, one weighted scorecard, and a mid-pilot call to cut the vendor that wasn't working.
A home care platform built with CAB, a caregiving company with 15+ years in the market who founded a software company, Sophicare, to build it with us. I was the only PM on a team of about 20 across UI/UX, DevOps, mobile, frontend and backend, over roughly two years.
Care plans drive the work: a client's required tasks get embedded in their profile and matched against visit hours. Open visits get filled by finding care partners whose skills fit and who live close enough to make the assignment realistic. Payroll and billing ran externally but pulled from an exportable report, and a dedicated mobile app let care partners clock in, clock out, and log completed tasks in the field.
Digitized the permitting process for construction on Montana's watersheds, with an applicant side and an admin side. Instead of a straight PDF-to-form port, we integrated Mapbox so reviewers could see exactly where construction was proposed and where existing permits sat, turning a stack of documents into decisions made with spatial context. A team of about eight developers plus UI/UX.
A structured conversation product in the family of video-call tools, but built around timed turns: a host gives each participant a fixed window to speak, say five minutes, then the floor passes automatically. The constraint is the feature, it forces the kind of turn-taking that free-for-all calls never produce.
Four startup discoveries. The core discipline holds across all of them, listen first, prioritize against what the business actually needs, map the user flows, and draw a defensible line between MVP and what waits for v2. What changes every time is the calibration: a phishing-defense product and a home-care platform don't get the same MVP, even when they get the same method.
Not everything is a two-year build. I've run three-week migration sprints, launched products from zero, and led projects the other direction, all the way to shutdown, owning the offboarding documentation so nothing got lost on the way out.
Outside client work, these are the products I've owned end to end. One is live, one closed, and all three taught me something the day job couldn't.
Natural hair care for kids, co-founded with a designer partner. Four SKUs, fully automated e-commerce, national fulfillment through Correos de Costa Rica. Built the operational backbone before touching a single Instagram post.
A study app built for a physician preparing for a physiatry specialty exam. Spaced repetition scored per topic, a two-voice audio pipeline for review on the move, and a driving-safe mode for the commute between hospital shifts.
An e-commerce brand I co-founded and ran for six years, alongside a full-time job. It worked while I could give it my weekends. What closed it wasn't a single bad call, it was a model built on assumptions that were true when we started and stopped being true over time: demand softened, and a business that leaned on my personal time to move orders couldn't absorb that. It's exactly why the next brand was designed so no part of it depended on me being physically available.
Before "Product Manager" was the title, this is where I learned to build the criteria a decision gets measured against, not just the decision itself.
Covid cost me my job the same year everything else got harder. So I did two things at once: I made empanadas through the early morning, sleeping in snatches until about 7, then went out to deliver them around 9, and I started Loggicare. Neither was a grand plan. One paid this month, the other was a bet on next year.
I'm not going to package it as a lesson. But running a business off my own weekends for six years, and watching it eventually close for exactly that reason, is why the brand I built after was designed so nothing depended on me being physically there.
Not a skills list. The criteria I use before a roadmap exists.
Most bad roadmaps aren't badly built, they're built on an assumption nobody stress-tested. I'd rather spend a week arguing about the premise than a quarter building the wrong solution well.
Given five options, the instinct is to rank them and pick one. Sometimes the actual answer is a routing rule: different tools for different technical constraints, not one tool for everyone.
A scorecard is only honest if the weights were set before you knew who'd win. Otherwise you're not evaluating, you're justifying.
When some inputs are more rigorous than others, the answer isn't to hide that. It's to name it, and fix the process next time so it's comparable from day one.
In one quarter the job can mean leading a 20-person build, owning a product from discovery to production, running a vendor evaluation, writing the system prompt, and translating all of it for executives. That's not scope creep, it's what the role actually is now. The title didn't shrink, the job description everyone keeps copying just hasn't caught up.
Moving fast is fine. Moving fast without a clear definition of done is just deferred rework.
A schedule slip, an underperforming vendor, a channel that isn't converting: none of it needs spin, and none of it needs alarm. It just needs to be said plainly and early.
The point of a pilot or an MVP isn't to ship something small, it's to find out whether the expensive version is even worth building, while the cost of being wrong is still low. Kill it early or scale it with evidence. Either one beats guessing.
A small model with just enough context about my work to answer honestly, including the parts that don't flatter me.
I like the ones that don't have an obvious answer yet.
A 90-day pilot to decide whether AI agents belonged inside a global fashion brand, and if that mattered, which platform to build on. The pilot started looking for one answer and ended with something more useful: a rule for deciding, per task, every time.
Agentic AI, systems that can do multi-step work rather than just answer questions, had crossed from novelty into something leadership could no longer ignore. The risk wasn't whether AI agents would show up inside the company. It was whether that would happen deliberately, or by accident, one department at a time, with no shared governance and no consistent security posture.
Five platforms claimed to solve this, spanning Microsoft-native builders, an external-connector builder, a self-contained reasoning model, and others. No objective framework existed to compare them across departments as different as Legal, Finance, HR, and Audit. There were no hard metrics to start from, so the case for action was strategic and risk-based, not a numbers pitch.
"Which platform do we buy" was the question everyone walked in with. It turned out to be the wrong one.
The implicit hypothesis going in: evaluate the top contenders rigorously enough, and one platform would emerge as the standard for the whole company. That's the natural framing when you're comparing five vendors against a scoring rubric, you're building toward a single winner.
The pilot didn't end that way. It ended with a routing model: different platforms for different technical situations, decided by what each tool could actually connect to, not by department preference.
This held up under pressure and got sharper by the end. The Phase 2 proposal extended it into an MCP Gateway concept, a layer that could enforce this routing consistently, plus a separate path for engineering work using native tools. The hypothesis didn't just survive contact with reality, it turned into an actual governance mechanism.
I was the sole Product Manager on the pilot. My piece: designing and owning the evaluation rubric, scoping the roughly 30 use cases across departments, weekly executive reporting, running enablement sessions with department AI champions, owning the issue tracker, and pulling five very different technical and strategic inputs into one coherent narrative for leadership, without losing the parts that were genuinely messy.
Worth naming directly: the platform that ranked first scored lowest of the field on governance and security. It won because build experience was weighted highest, and that's where most of the real task volume lived. A weighted scorecard doesn't produce the tool that's best at everything, it produces the tool that's best at what you decided mattered. That's a feature of the method, not a flaw in the winner.
Department names aren't the point here, the pattern is: readiness split cleanly into tiers based on how much process rigor a function already had, not how excited they were about AI.
No native file generation in workflow, unreliable enterprise integration, and stability problems severe enough to crash live demos. That's a clean, documentable failure, and it's why it was cut before the finish line.
Some champions deeply customized their tools. Others barely engaged, and some platforms went unassessed simply because nobody explored them. That means part of the evaluation data reflects how much effort a person put in, not just how good the tool was.
Identical prompts sometimes produced different formatting across runs, and large pasted context visibly degraded reasoning in at least one platform. Worth naming so the ranking doesn't read as more certain than it is.
Most issues were escalated to the vendors responsible rather than resolved internally, which is the right call when the fix isn't yours to make. A meaningful share traced back to the builder that was ultimately discontinued.
Standardize champion time commitment before scoring starts, a minimum number of hours or sessions required per champion, so the evaluation data is comparable across platforms from day one instead of discovering the unevenness after the fact.
The right output of a pilot isn't a single winner, it's a decision framework. Executives walk in wanting to know which tool to buy. The more useful deliverable was the routing logic itself, a reusable way to make that call per use case, indefinitely, instead of a one-time pick that goes stale the moment the tools evolve.