Hang Yao — Data Scientist

Hang YaoHang Yao姚航

data scientist · data engineer · barcelona data scientist · data engineer · barcelona 数据科学家 · 数据工程师 · 巴塞罗那

● available immediately · Barcelona · open to relocate● disponible de inmediato · Barcelona · abierto a reubicarme● 2026届应届硕士 · 随时可到岗 · 巴塞罗那

profileperfil个人总结

Results-driven data professional across the full data lifecycle — from scalable architectures and ETL pipelines to predictive modeling with foundation models (LLMs/VLMs) and applied analytics. I design end-to-end data solutions and take complex ML workflows to production with Spark, Airflow, DuckDB and PyTorch, backed by enterprise software engineering experience at NTT DATA.

Profesional de datos orientado a resultados en todo el ciclo de vida del dato — de arquitecturas escalables y pipelines ETL al modelado predictivo con modelos fundacionales (LLMs/VLMs) y analítica aplicada. Diseño soluciones de datos de extremo a extremo y llevo flujos de ML complejos a producción con Spark, Airflow, DuckDB y PyTorch, respaldado por experiencia en NTT DATA.

结果导向的数据全栈人才,覆盖数据全生命周期——从可扩展架构与 ETL 流水线,到基于基础大模型 (LLMs/VLMs) 的预测建模与应用分析。使用 Spark、Airflow、DuckDB 和 PyTorch 设计端到端数据方案并将复杂 ML 工作流推向生产,辅以 NTT DATA 的企业级软件工程经验。

90.2%structurally valid VLM outputssalidas VLM válidasVLM 结构化有效输出
150k+API records modeledregistros de API modelados建模的 API 记录
10 GBheterogeneous data integrateddatos integrados集成的多源数据
~90%test coverage sustainedcobertura de tests测试覆盖率

nowBuilding a project on experimentation — A/B analysis + dbt modeling + an LLM agent.Construyendo un proyecto de experimentación — análisis A/B + modelado con dbt + un agente LLM.正在做一个实验分析项目——A/B 分析 + dbt 数据建模 + LLM agent。

experienceexperiencia工作经验
2024.02 – 2025.06

Back-End Engineer InternIngeniero Back-End (Prácticas)后端开发工程师(实习)

NTT DATA · Barcelona · NTT Group (Fortune Global 500)Barcelona · Grupo NTT (Fortune Global 500)巴塞罗那 · 世界500强 NTT 集团旗下

  • Built enterprise-grade microservices and state-driven workflows on Spring to securely route and process financial records for a major banking client.
  • Desarrollé microservicios de nivel empresarial y flujos basados en estados con Spring para procesar de forma segura registros financieros de un gran cliente bancario.
  • 使用 Spring 生态开发企业级微服务与状态驱动的工作流,为大型银行客户安全地路由和处理金融记录。
  • Mapped UI states to custom relational schemas with advanced SQL integrations for fast, reliable data persistence.
  • Mapeé estados de interfaz a esquemas relacionales a medida con integraciones SQL avanzadas para una persistencia rápida y fiable.
  • 将 UI 状态映射至定制关系型数据库架构,执行高级 SQL 集成以确保高效可靠的数据持久化。
  • Orchestrated end-to-end delivery via GitHub Actions, resolving production bugs and sustaining ~90% unit-test coverage with JUnit.
  • Orquesté la entrega con GitHub Actions, resolviendo bugs de producción y manteniendo ~90% de cobertura de tests con JUnit.
  • 通过 GitHub Actions 编排端到端交付,解决生产环境 Bug,并用 JUnit 将核心模块测试覆盖率维持在约 90%。
JavaSpring BootSQLJUnitGitHub Actions
projectsproyectos项目经历
vlm · thesis

Floor Plan Graph Extraction with Vision-Language ModelsExtracción de Grafos desde Planos con VLMs基于视觉语言模型 (VLM) 的户型图拓扑提取

  • Annotation-free pipeline converting 1,276 raster floor plans into structured topological graphs with a frozen VLM (Qwen3-VL-8B) — no fine-tuning, no labels.
  • Pipeline sin anotaciones que convierte 1.276 planos ráster en grafos topológicos con un VLM congelado (Qwen3-VL-8B), sin fine-tuning ni etiquetas.
  • 免标注流水线,用冻结的 VLM (Qwen3-VL-8B) 将 1,276 张户型图转为结构化拓扑图,无需微调或标注。
  • Six-stage prompting strategy reaching 90% classification accuracy and structurally valid graphs for 90.2% of plans; built an automated validation framework and human-in-the-loop refinement.
  • Estrategia de prompting en seis etapas con 90% de precisión y grafos válidos en el 90,2% de los planos; marco de validación automático y refinamiento human-in-the-loop.
  • 六阶段提示词策略达 90% 分类准确率,90.2% 户型图结构有效;构建了自动验证框架与人机协同修正系统。
Qwen3-VLVLMsPrompt EngineeringPyTorch
imageraster plan
1 · classifyaggr / indiv / other
2 · 2D graphnodes + edges
3 · enrichroom counts
4 · validatechecks + flags
6 · stack 3DN floors

side-loop (--review): [4] flags → [5] human-in-the-loop hints → corrected graph → re-run [4]+[6]

gen-ai

STGF-LLM: Time-Series Forecasting & Synthetic Generation with LLMsSTGF-LLM: Series Temporales y Generación Sintética con LLMsSTGF-LLM:基于 LLM 的时序预测与合成数据生成

  • Fine-tuned an LLM (Qwen3) for robust time-series forecasting and high-fidelity synthetic data generation, with a history-aware prompting strategy to serialize multivariate temporal data.
  • Fine-tuning de un LLM (Qwen3) para previsión de series temporales y generación sintética de alta fidelidad, con prompting "history-aware" para datos temporales multivariantes.
  • 微调 LLM (Qwen3) 实现稳健的时序预测与高保真合成数据生成,设计"历史感知"提示策略序列化多变量时序数据。
  • Benchmarked against TimeGAN and Chronos-T5 on Train-on-Synthetic-Test-on-Real accuracy, competitive within hardware constraints.
  • Comparado con TimeGAN y Chronos-T5 en precisión TSTR, competitivo dentro de las restricciones de hardware.
  • 与 TimeGAN、Chronos-T5 在 TSTR 准确率上基准测试,硬件受限下仍具竞争力。
Qwen3LLMsGenerative AITime-Series
STGF-LLM two-phase architecture: instruction-tuned LoRA training and autoregressive sliding-window inference
architecture — phase 1: semantic serialization + LoRA fine-tuning · phase 2: autoregressive sliding-window generation
mlops

Steam Games Satisfaction Prediction — MLOps PipelinePredicción de Satisfacción en Steam — MLOpsSteam 游戏满意度预测 — MLOps 流水线

  • Full ML pipeline on 150,000+ Steam API records: profiling, PCA feature selection, and a custom neural network to maximize predictive R².
  • Pipeline ML completo sobre 150.000+ registros de Steam: profiling, selección de variables con PCA y una red neuronal a medida para maximizar el R².
  • 在 15 万+ Steam API 记录上的完整 ML 流水线:数据分析、PCA 特征选择、定制神经网络最大化预测 R²。
  • MLOps pipeline with DuckDB for full data lineage and instant rollback, plus a custom UI for automated execution and monitoring.
  • Pipeline MLOps con DuckDB para linaje completo y rollback instantáneo, con UI para ejecución y monitorización.
  • MLOps 流水线用 DuckDB 保证完整数据血缘与即时回滚,配定制 UI 实现自动执行与监控。
PyTorchPCADuckDBMLOps
Steam API150k+ games
formatted.duckdb
trusted.duckdb
exploitation.duckdb
featuresPCA
predictneural net · R²

each zone = separate DuckDB instance → full data lineage + instant rollback

lakehouse

Big Data Management — End-to-End Data LakehouseGestión de Big Data — Data Lakehouse大数据管理 — 端到端数据湖仓

  • Containerized 3-tier (Cold/Warm/Hot) lakehouse on Docker, integrating 10 GB of heterogeneous sources via REST APIs with Spark and Redis; daily ETL orchestrated with Airflow.
  • Data lakehouse de 3 niveles en Docker, integrando 10 GB de fuentes vía REST APIs con Spark y Redis; ETL diario orquestado con Airflow.
  • Docker 上的三层(冷/温/热)数据湖仓,用 Spark 和 Redis 通过 REST API 集成 10GB 多源数据,Airflow 编排日常 ETL。
SparkAirflowDockerDelta Lake
Big Data lakehouse architecture: Landing, Trusted and Exploitation zones on Delta Lake, with cold/hot/warm ingestion paths orchestrated by Airflow
architecture — cold/hot/warm ingestion · Landing → Trusted → Exploitation zones on Delta Lake · Airflow + Docker
educationeducación教育背景
2025.09 – 2026.01

Exchange Student (Master's)Estudiante de Intercambio (Máster)硕士交换生86/100

Beihang University · Beijing · China's top-10 in engineeringUniversidad de Beihang · Pekín · top-10 de China en ingeniería北京航空航天大学 · 北京

2024.09 – 2026.06

MSc Data ScienceMáster en Data Science数据科学硕士7.6/10 · ES scale · 西班牙制

Universitat Politècnica de Catalunya (UPC) · QS Eng. & Tech. #96 · QS 工程与技术全球第96

2020.09 – 2024.06

BSc Computer ScienceGrado en Ingeniería Informática计算机科学学士7.3/10

Universitat Politècnica de Catalunya (UPC)

skills & languageshabilidades e idiomas技能与语言
data enging. datos数据工程
SparkAirflowDockerDuckDBETLMLOps
ml / aiml / ia机器学习
PyTorchTensorFlowScikit-LearnLLMsVLMsPCA
languagesidiomas编程
PythonSQLJavaGitREST APIs
spokenidiomas语言
Chinese · nativeChino · nativo中文 · 母语English · fluentInglés · fluido英语 · 流利Spanish · fluentEspañol · fluido西班牙语 · 流利Catalan · fluentCatalán · fluido加泰罗尼亚语 · 流利
© 2026 Hang Yao · built as an AI-native page — query me on the right →página AI-native — pregúntame a la derecha →AI-native 页面 — 在右侧向我提问 →