Data Science × Applied AI BNP Paribas Data Office · 2026

Rigorous models.
Visible evidence.

I am Ibrahim Abdelatif, an M2 Data Science student at Paris-Dauphine and Agentic AI Data Scientist Intern at BNP Paribas. I build statistical work that can be challenged and AI systems whose boundaries can be inspected.

4 flagship systems

Case studies that show the decision trail.

Each case exposes the problem, architecture, validation scope, result and limitation. A recruiter should not need to guess what was actually built.

01 / Hardened research demo

Gemma 4 Hackathon · 2026

ImciFlow

LLM extraction, RAG evidence and deterministic IMCI rules in one auditable workflow.

A multilingual clinical decision-support prototype that combines Gemma 4, retrieval over IMCI references, deterministic safety rules and a human-review boundary.

3

languages

English, French and Sudanese Arabic paths

Gemma 4LangGraphFastAPIReactChroma
The evaluation does not measure real Gemma extraction quality, RAG relevance, hallucinations or clinical outcomes. Shared authentication, distributed rate limiting and encrypted durable storage remain production work. The demo must not be used for diagnosis.
Evidence case study
01020304PROFILE → RECOMMEND → VALIDATE → APPROVE
02 / Tested prototype

Applied AI engineering · 2026

GenAI Data Prep

Deterministic data-quality checks first; LLM recommendations second.

A LangGraph workflow that profiles a dataset, proposes preprocessing decisions, validates transformations and keeps a human approval point before output generation.

84/84

tests

Deterministic tests run without a provider key

LangGraphFastAPIStreamlitPydanticOpenAI
The hardened default does not send raw rows and constrains local paths and URLs, but production use still requires explicit consent, secret governance, encrypted retention and provider monitoring.
Evidence case study
03 / Tested synthetic decision lab

Independent ML project · 2025

Credit Risk

Chronological validation, calibrated probabilities, decision cost and subgroup diagnostics.

A synthetic approval-decision laboratory that separates fit, calibration and final holdout periods, selects a cost-sensitive threshold and exposes subgroup diagnostics through reproducible reports and an API.

0.9669

ROC-AUC

Final chronological holdout of 4,000 synthetic rows

Pythonscikit-learnCalibrationFairness diagnosticsFastAPI
The data and approval mechanism are synthetic, so the system does not estimate real default risk. A 14.3-point age-band selection-rate spread is disclosed as a risk requiring contextual investigation, not proof of fairness.
Evidence case study
04 / Live demo · Public source

Paris-Dauphine · 2025

Segmentation

From clustering diagnostics to an interface a marketing team can actually explore.

An R/Shiny product that compares unsupervised methods, profiles four customer groups and makes the assumptions and cluster diagnostics inspectable.

2,240

customers

Behavioral and demographic observations

RShinyK-meansCAHGMM
The segments have not yet been validated through campaign uplift, temporal stability or out-of-sample assignment. They support exploration, not causal targeting claims.
Evidence case study

How I decide whether a project is ready to show.

01

Evidence before adjectives

Metrics include their evaluation scope. Deployed means deployed; a prototype stays a prototype.

02

Deterministic where it matters

LLMs recommend and structure. Rules, contracts and human review protect high-impact decisions.

03

A model needs an interface

The work is not finished when a notebook ends. APIs, tests and usable screens make it inspectable.

Current Agentic AI internship at BNP Paribas, prior forecasting and reporting work at Deloitte, and a mathematical path through Strasbourg and Paris-Dauphine.

Mathematics is the foundation. Delivery is the test.

  1. Apr — Sep 2026

    01 / Experience

    Agentic AI Data Scientist Intern

    BNP Paribas — Data Office Europe Mediterranean

    Agentic AI experimentation, internal LLM platform workflows, use-case framing and evaluation across business, IT and data constraints.

  2. Jun — Sep 2025

    02 / Experience

    Data Scientist Intern

    Deloitte Chad

    Time-series budgeting models, reporting automation and translation of finance needs into analytical specifications.

  3. Oct 2024 — May 2025

    03 / Research

    Junior Researcher — Non-linear dimension reduction

    Paris-Dauphine University

    Comparative PCA/KPCA analysis and a Python workflow for latent-space visualisation.

  4. 2025 — 2026

    04 / Education

    M2 Statistical & Financial Engineering — Data Science

    Paris-Dauphine University — PSL

    Deep learning, NLP, reinforcement learning, cybersecurity, data quality and climate-risk modelling.

  5. 2024 — 2025

    05 / Education

    M1 Applied Mathematics — Statistics

    Paris-Dauphine University — PSL

    Statistical learning, GLMs, stochastic processes, optimisation and scientific computing.

  6. 2021 — 2024

    06 / Education

    BSc Applied Mathematics

    University of Strasbourg

    Mathematics, algorithms, modelling and scientific computing foundations.

A stack organised by what it enables.

A

Model with evidence

Statistical foundations, explicit baselines and evaluation protocols before headline metrics.

scikit-learnXGBoostTime seriesClusteringMonte CarloEconometrics
6 curated case studies
B

Engineer the boundary

APIs, structured outputs and deterministic safeguards around model behavior.

FastAPIPydanticLangGraphRAGDockerSQL
81 tests verified
C

Ship for a user

Interfaces and deployments that make assumptions, evidence and limitations inspectable.

ReactNext.jsR ShinyCloud RunVercelGit
2 live products

Have a hard data problem? Let’s make the evidence visible.

I am interested in Data Science and Applied AI work where model quality, software reliability and stakeholder clarity matter together.

Signal available

Paris · Europe · French / English / Arabic