AI Engineer
2026Outlier AI · Agent Evaluation / OpenClaw
Contributed to the rigorous evaluation of frontier multimodal agents on OpenClaw by engineering adversarial, production-grade user scenarios across heterogeneous data ecosystems—including users, messages, emails, contacts, FinTrack, Strava, Notion, and many other sources. Acting as the end user, I crafted deceptively simple prompts that forced agents to resolve ambiguity, cross-reference conflicting signals, and fuse structured database state with multimodal evidence (PDFs, images, videos, and audio) under strict correctness constraints. I audited every requirement and deliverable with fine-grained rubrics across models such as Claude Opus 5, OpenAI GPT-5, and Gemini 3.1 Pro, isolating failure modes and producing high-signal datasets of successful and failed trajectories that clients use to retrain and harden models on OpenClaw.
AI AgentsMCPJSON handlingAPIsDatabasesCode execution environmentsCSV handling
AI Engineer
2026Outlier AI · Agent Evaluation & MCP Tooling
Evaluated and enhanced autonomous AI agents operating within tool-integrated environments (MCP systems), including file systems, Git, databases, mapping tools, and code execution frameworks. I designed adversarial and high-difficulty scenarios to identify reasoning breakdowns, trajectory failures, and tool misuse. For each failure, I engineered the ideal execution pathway, corrected outputs, and developed structured rubrics to systematically improve performance. This work significantly increased agent robustness, task completion reliability, and multi-tool reasoning coherence.
AI AgentsPythonMCPLinuxJSON handlingAPIsdatabasescode execution environments
Data Scientist
2025 - 2026Outlier AI · LLM Evaluation for Data Science Workflows
Contributed to the improvement of LLMs in data science workflows. I designed complex tasks involving large-scale datasets that required precise coding, statistical operations, and high-fidelity visualizations. After prompting LLMs to generate solutions, I evaluated their outputs using structured rubrics, identifying logical gaps, implementation errors, and deviations from constraints. When models failed, I produced an ideal, step-by-step reasoning and implementation path the model should follow to get the correct response (numerical outputs and visualizations), ensuring strict adherence to requirements.
PythonpandasnumpymatplotlibseabornscipysklearnLLM promptingrubric-based evaluationdata visualizationGoogle ColabJupyter NotebooksCSV handlingJSON handlingXLSX handling
AI Engineer
2025Outlier AI · Autonomous Build & Debugging Agents
Contributed to the refinement of AI agents responsible for diagnosing failed software builds and broken repositories. The agents analyzed execution logs, dependency errors, and test failures to determine root causes and implement corrective patches. When the agent's reasoning diverged from best practices or when failed to solve the problem, I reconstructed the correct debugging trajectory, improved scripts, reinforced unit and custom test validation, and ensured successful rebuilds. This role required deep understanding of CI/CD workflows, repository architecture, and automated debugging methodologies.
AI AgentsNode.jsCLICI/CDDockerGitLinuxGitHubpythonJavascriptKotlinother languagestesting frameworks
Software Engineer
2025Outlier AI · LLM Code Generation Refinement
Focused on improving LLM-generated code through iterative refinement cycles. I evaluated model outputs against strict instruction sets, performance constraints, efficiency requirements, and formatting standards. When deficiencies were detected, I manually refactored and optimized logic, corrected edge cases, and enhanced computational performance. This multi-turn feedback loop strengthened the model’s ability to produce clean, efficient, and instruction-compliant code, driving outputs toward production-level reliability.
Pythoncode reviewoptimizationpandasnumpyseabornmatplotlibscipysklearnother libraries
Software Engineer
2025Outlier AI · Autonomous Code Repair Systems
Contributed to the development of autonomous code-repair agents by generating high-quality repair trajectories across real-world GitHub repositories. I manually simulated the agent’s end-to-end reasoning process: analyzing issue reports, inspecting repository architecture, reproducing failures, identifying root causes, designing compliant patches, and validating fixes through rigorous unit and custom test construction. These structured trajectories served as supervised training data to teach the agent how to follow strict software engineering best practices, including clean code principles, robust test coverage, and reproducible Linux-based execution workflows.
PythonGitHubpytestLinux command linecode reviewdebuggingtestingpatchingpandasnumpyother libraries
Data Scientist
2025Outlier AI · Spreadsheet & Structured Data Intelligence
Contributed to the evaluation and correction of LLMs designed to generate spreadsheet-based solutions from large, structured datasets. My role involved stress-testing models against detailed data requirements, auditing formula accuracy, logical consistency, and compliance with formatting constraints. When models failed, I reconstructed the correct analytical pathway, delivered the accurate solution, and documented failure modes. Using granular evaluation rubrics, I ensured every output met strict standards of numerical precision and structural correctness, reinforcing model reliability in enterprise-grade data workflows.
PythonpandasnumpyopenpyxlLLM promptingrubric-based evaluationdata analysisCSV handlingJSON handlingXLSX handling
Mathematician
2024 - 2025Outlier AI · LLM Mathematical Reasoning Specialist
Improved mathematical reasoning capabilities of Large Language Models across diverse domains, including algebra, calculus, geometry, and advanced problem-solving. I designed complex natural prompts to stress-test logical rigor and symbolic correctness. When models produced incorrect or incomplete solutions, I reconstructed fully rigorous mathematical derivations that strictly followed formal principles and constraints. My contributions strengthened symbolic accuracy, step-by-step reasoning integrity, and compliance with mathematical standards.
LaTeXGeogebraSymbolabalgebracalculusgeometryarithmeticnumber theorycombinatoricsother mathematical domains
Full Stack Developer
ConfidentialPrivate Enterprise Clients
Designed and delivered end-to-end web applications integrating full-stack development with embedded data science components. I architected scalable backend systems optimized for high-throughput data processing, engineered efficient database querying strategies, and built interactive, analytics-driven frontend interfaces. Additionally, I implemented containerized environments and CI/CD pipelines to ensure reproducibility, deployment stability, and infrastructure consistency. All projects were developed under strict confidentiality agreements.
PythonJavaScriptReactNode.jsSQLDockerdockerGitHubSQLAlchemypandasnumpyCSV handlingJSON handlingGit
SOCIALMEDIA
LET'S CONNECT
SOCIAL MEDIA
LET'S CONNECT
SOCIAL MEDIA
Alvaro Vasquez
@alvarovasquez.ai
My Objective
We will live in a world transformed by AI, and knowledge is our greatest tool. Often, the theory and formulas are a barrier. My objective is to tear it down using intuitive animations, allowing anyone to understand the fundamentals of AI from scratch to adapt and thrive in this new era.
Featured Content