Subash Ale Magar

Data Engineer · AI & Data Architecture

Cologne, Germany

Building the systems behind useful AI.

I’m Subash, a Data Engineer with a background in data science. I own the architecture, infrastructure, and implementation of semantic search, and build the data and model services that AI applications depend on.

Search a small universe of ideas

Semantic search groups ideas by meaning. Search to see a neighbourhood light up.

This interactive demo needs JavaScript.

Search or pick an example. The eight nearest neighbours appear here, closest first.

An interactive illustration of semantic search, using synthetic data and simulated matching.

Currently at LexisNexisSemantic searchAgentic applicationsModel serving & MLOps

Selected work

What I’m building.

Recent work at LexisNexis, from search services to AI applications.

Owned end to end

Semantic search for patent data

Connecting patent information through meaning. I’m responsible for the service from vector ingestion and search APIs through to deployment and day-to-day operations.

Python / FastAPIQdrantDatabricksAWS / Kubernetes
Problem & approach

Problem. Patent terminology varies. Applications need a reusable way to retrieve related information by meaning, alongside existing search approaches.

Design. Separate model inference, vector ingestion, and search APIs so that consumers can reuse the capabilities and components can evolve independently.

Operations. Deployment, access controls, observability, and keeping the service reliable as usage grows.

Applications

Agentic AI applications

I contribute to the delivery of agentic applications, connecting language models with tools, retrieval, and enterprise data.

Problem & approach

Problem. An agent needs access to the right information and tools to complete a useful workflow. Model responses alone are only part of the application.

Contribution. I help deliver the application by connecting data, retrieval, and AI services with user-facing capabilities.

Focus. Reusable service interfaces and clear boundaries between the agent, its tools, and underlying data systems.

Model operations

Inference on Databricks

I build model inference capabilities on Databricks, bringing embedding models into application and data workflows with MLOps.

Problem & approach

Problem. Embedding models must support both data processing and application requests, with repeatable deployment and evaluation.

Implementation. I build inference capabilities on Databricks and connect them with embedding pipelines and downstream services.

Operations. My work includes MLOps processes for evaluating and maintaining model-based capabilities.

Knowledge & retrieval

RAG & knowledge graphs

My work spans retrieval-augmented generation and knowledge graphs, connecting AI applications to patent information and its relationships.

Problem & approach

Problem. Similarity retrieval finds related content, while some questions depend on explicit entities, relationships, and source evidence.

Approach. I work on combining retrieval with structured knowledge, with interests in patent claims, data linkage, and AI access to enterprise information.

Scope. This is ongoing work; I distinguish research and prototypes from delivered application capabilities.

Experience

A foundation across disciplines.

Software development, data science, and now data engineering.

Feb 2023 — PresentBonn, Germany

LexisNexis

Intellectual Property Solutions

Data Engineer · Previously Data Scientist III

Responsible for the semantic search service end to end. Contribute to agentic application delivery, Databricks model inference, and work on RAG and knowledge graphs. Earlier work includes patent matching, graph-based patent family creation, and custom classification.

Apr 2021 — Jan 2023Cologne, Germany

nextmarkets

Data Scientist

Built and maintained lakehouse platforms using Azure and Databricks. Developed data pipelines, risk assessment automation, client scoring, fraud detection, and analytics for internal teams.

Apr 2019 — Mar 2021Passau, Germany

University of Passau

Research Assistant & Developer

Developed interactive research applications and OCR workflows. Supported teaching in information retrieval, database systems, and web science.

Earlier experienceNepal & Germany

Software development & systems

Web Development · System Administration

Built websites and business applications using PHP, JavaScript, Java, and databases. Worked in system administration and support, establishing a practical foundation in software and infrastructure.

About me

Curious about models.
Responsible for systems.

My career has grown from software development and systems administration into data science and data engineering. I bring those perspectives together to build AI capabilities that fit the systems around them.

I enjoy taking ownership of complex problems: understanding the need, designing the architecture, implementing the solution, and supporting its operation.

Originally from Nepal and now based in Cologne, I hold a master’s degree in Computer Science from the University of Passau, with a focus on data science, machine learning, and text mining.

Education

MSc Computer Science · University of Passau
BSc Computer Science · Pokhara University

Community

My volunteer work has included website development for Apnups in Uganda and humanitarian initiatives with Nyayik Sansar in Nepal.

Data & infrastructure

Python · SQL · PySpark · Databricks
Azure · AWS · Data pipelines

AI & retrieval

Machine learning · Model serving · MLOps
Semantic search · RAG · Knowledge graphs

Services & delivery

FastAPI · Vector databases · Kubernetes
Architecture · CI/CD · Observability

Contact

Let’s connect.

Have a question about my work, a technical challenge, or an idea to exchange? Get in touch.