Software engineer · Ankara, Türkiye

Salih Karagöllü

I build backend services at TÜBİTAK BİLGEM for Türkiye’s national public financial management system, which has more than 2 million users. Java and Spring services that talk over Kafka, PostgreSQL underneath, React on top. Most of the interesting problems are about keeping state correct when several services touch the same record.

I also do applied computer vision. My graduation project became a first-author IEEE paper on dental segmentation that runs live on an Android phone.

Work

Problems I’ve worked on at TÜBİTAK BİLGEM, from cross-service workflows to production recovery and the tools other teams now use.

Distributed workflows

Approvals that can’t apply to data the user never saw

Two services, a Kafka topic between them, and a record that can change while an approval is in flight. I designed the flow and wrote the guide other developers follow to add new approval steps.

Travel expense claims live in two services. Accounting officers edit a claim in the service that owns it. Supervisors approve it from a second service, which sends the approval over Kafka.

Checking that the claim is still awaiting approval isn’t enough. A supervisor can open a claim for ₺5,000; meanwhile the officer pulls it back, changes it to ₺50,000 and resubmits it. When the supervisor clicks approve on the old screen, the claim is awaiting approval again. Same status, different content: the ABA problem.

I designed the flow around two rules: the owning service is the only one that decides, and every action carries the version the user was looking at.

  • The record’s version goes from the supervisor’s screen through the controller and mapping layer into the Kafka command. The owning service compares it with the current @Version before changing anything, and rejects a mismatch.
  • The requesting service may prepare documents or an approval date, but it doesn’t write them as final until the owner answers. Success and failure are separate result messages, and a failure can’t carry changes to apply.
  • On failure, the owner deletes the documents the action generated, because a database rollback doesn’t remove them.
  • Edits inside the owning service stay synchronous and use the same version check. Only actions that cross the service boundary go through Kafka.
  • Each service ignores events it published itself on the shared topic, and a result for a record deleted in the meantime is logged and dropped instead of failing the consumer on every retry.
  1. Supervisor’s screen: Opens the claimv1 · ₺5,000
  2. Owning service: Officer edits and resubmitsv4 · ₺50,000
  3. Screen to requesting service: approve, v1
  4. Requesting service to owning service: command, expects v1 Kafka
  5. Owning service: v1 ≠ v4nothing applied
  6. Owning service to requesting service: result: failed Kafka
  7. Requesting service to screen: claim has changed
What happens to the approval from the old screen.

What it costs

Supervisors get the outcome asynchronously, so the UI confirms the action was received and reports the result later. When an officer and a supervisor act at the same moment, the officer’s synchronous write usually wins; the supervisor gets a “claim has changed” message, never a corrupted claim.

I wrote the design up with the alternatives we rejected, plus an implementation guide that other developers use to add new approval steps.

Read the design note

Developer tooling

Tools other teams use

Built for my team, now used by several others. What each one changed:

  • Release extension 21% → 1% of a developer’s week spent on releases
  • Backport skill 33% → 8% of a fix’s time spent on backports
  • Review agent N+1 queries caught before merge
  • Data masking External LLMs usable on sensitive data

I built these for my own team first. Several other teams have since adopted them, and I wrote the setup and usage guides for each.

Share of time, before and after
Releasesof a developer’s week before 21% after ~1%
Backportsof a fix’s time before 33% after ~8%

Release extension: from a fifth of the week to almost nothing

Releases used to take at least 21% of a developer’s week on average. Now they take about 1%. A release means tagging issues, cutting branches in the right order, calculating versions for shared libraries and the applications that consume them, writing release notes and starting builds, across several repositories at once. The browser extension works out what a release needs and does all of it, for one repository or many. It used to apply database changelogs and sync deployments too; I took that out so a developer stays in control of production changes.

Backport skill: from a third of a fix’s time to 8%

Getting a fix onto every supported release line means branches, pull requests, cherry-picks, matching library and application versions, and release notes for each line. That used to eat about a third of the time spent on a fix. An agent skill now plans it and opens the pull requests, and the share is down to about 8%.

Review agent: N+1 queries stopped slipping through

Our human reviews kept missing problems like N+1 queries. The review command and subagent check a pull request, a set of branches or local changes in five to ten minutes, and that class of issue now gets caught before merge.

Data masking: external LLMs became usable at all

Our logs and database rows contain personal data, so nobody could paste them into an external AI tool. A macOS menu-bar app masks the sensitive fields on the clipboard first. What leaves the machine is anonymized, and the team can use external LLMs for debugging.

Data

One query, too many rows

An expenditure’s detail view loaded several independent document collections in a single query. Joining collections that don’t depend on each other returns one row for every combination, so the result grows with the product of their sizes, not the sum. Some expenditures have hundreds of related documents.

Three collection joins in one query 20 × 30 × 10 = 6,000 rows
One targeted query per collection 20 + 30 + 10 = 60 rows

I split the fetch into one query per collection and assembled the object graph in a helper, with a separate path for callers that also need nested documents. The persistence context maps every result back onto the same entities, so callers still get one complete graph.

  • In batch payment validation I fixed the opposite problem: one query per payment order inside a loop. A fetch specification loads them with their parent, and a local check skips the remote validation call when there is nothing to validate.
  • In a paginated list, putting the fetch on the pageable query broke paging, and an inner join hid records without an optional relation. Filtering and paging now stay in the base query, relations load separately, and the optional join is a left join.
  • Repository tests flush and clear the persistence context before asserting with Hibernate.isInitialized. Otherwise entities left over from the test setup look loaded and a missing fetch passes.

Production

Recovering lost callbacks without touching the main consumer

A callback from another service occasionally never arrived, which left an employee’s record waiting indefinitely. While the underlying cause was being investigated, I added a recovery job.

  • It runs on a single instance and can be switched off by a parameter.
  • It uses its own Kafka consumer outside the main group, without auto-commit, so the main consumer’s offsets never move.
  • It assigns partitions directly and seeks by timestamp to scan a bounded window, from two hours to thirty minutes before each run, so it never races fresh messages.
  • Each event goes through the existing linking logic, which does nothing when the record is already linked and never overwrites a different link.

It contains the problem; it isn’t the fix for the root cause, and isn’t meant to be.

Production

Timeouts that looked like a Kafka problem

Some services logged Kafka poll timeouts, lost group membership and failed offset commits. It looked like a messaging problem. Reading the logs in order showed database read timeouts during updates first: the consumer was blocked on the database long enough to be removed from its group, and the Kafka errors followed.

I worked through the remaining hypotheses with infrastructure and network colleagues: row locks, dropped connections, a mismatched build, oversized updates. Two things I had to keep straight: missing rows during a long first attempt don’t prove a message was lost, and a read replica can’t tell you what happened on the primary. I also reverted a speculative entity change so the network team could reproduce the original behavior cleanly.

Production

Which lost events are still safe to replay?

A set of historical relationship events between two services had apparently never been processed. Replaying all of them would have been easy and wrong: since then, some records had been relinked, cancelled or paid, and an old event could undo a valid current state.

I worked out the side effects of each event type and wrote a read-only SQL classification over message content, consume times, current relations, cancellations and payment history, including possible message-ID collisions. Each of the 33 events got a verdict someone else could check before anything was republished.

Debugging

A browser bug that was really a deployment path

Compressing attachments in the browser worked locally and hung on every deployed environment. I compared the conversion code, browser profiles, library versions, callback order and memory copies before finding the actual cause.

A script the compression library loads at runtime was served from the wrong path in the Nginx configuration. The local dev server served it correctly, which hid the problem. The fix was one line in each of three environments.

Distributed workflows

Commit first, acknowledge second

A shared library acknowledges a Kafka message once the listener method returns. With nested transaction interception, “returned” can come before the outer transaction physically commits, so a message can be acknowledged while its data is still at risk.

I traced the Spring proxy and advice order and wrote reproduction tests for the rules the library has to follow: acknowledge only after the physical commit, never when the commit fails, and roll back before a negative acknowledgment.

Integrations

Turkish characters across two hops

A name search goes through a shared library and an outbound gateway before reaching an external service, and Turkish characters broke along the way. There were two separate problems: the query string wasn’t explicitly encoded as UTF-8, and the gateway built its JSON body as a string instead of serializing a typed request.

I fixed both and changed the contract tests to use Turkish characters. ASCII fixtures pass either way, which is how the bug got through.

Security

Having the role isn’t having access

A lookup by petition number checked that the user had the right role, but not that the record belonged to the user’s own public administration. In a system shared by many institutions, that difference matters.

I added a record-level check that loads the record’s administration and compares it with the user’s, and routed the alternate lookup paths through the same check. Separately, I narrowed which roles can see preliminary control opinions before a document is formally sent.

Integrations

Tax-debt petitions across two services

Payment orders need to reference tax-debt petitions that another service owns. I built the feature across both services and the client.

  • A relationship table with Liquibase migrations.
  • Secured service-to-service endpoints and query rules for eligible petitions: calculated, matching taxpayer identity, inside a 15-day validity window.
  • A React search-and-select screen.
  • An audit trail of every addition, update and removal.
  • Adding or removing a petition is asynchronous, so the payment side keeps a pending state until the other service confirms.

Data

Payment types moved from code to data

Payment types were hard-coded as Java enums and JavaScript constants, so every regulatory change meant a deployment across several repositories.

I moved them into a table, migrated more than 150 existing records from string values to foreign keys, exposed the types through an endpoint, injected the lookup into the mapping layer and replaced the client’s static enum with a cached context. New payment types have since been added without a deployment.

A follow-up fixed lowercase codes arriving from upstream integrations with a Turkish-locale uppercase conversion, since a lowercase i doesn’t uppercase to I in Turkish.

Research

Computer vision, from the dataset to a phone.

DentAI

First author, IEEE ASYU 2025

Read the paper

A graduation project at Gebze Technical University, supported by HAVELSAN’s AI JET BİGG program. I started with one model for everything. The data made that a bad idea: clinical gingiva photos are standardized, consumer photos of teeth are not, and the two uses need different trade-offs. So there are two models: a gingiva model for clinical analysis, and a DeepLabV3-MobileNetV3 that runs on the phone.

Intraoral photo of an upper dental arch with several fillings.
Photo
Annotated ground-truth mask for the photo.
Ground truth
The model’s predicted mask for the photo.
Prediction
Mean pixel accuracy
95.79%
Weighted IoU
90.10%
Dataset
2,495 images

Mobile model on its test split: teeth, caries, cavities and cracks in phone-quality photos.

The gingiva model against the published result

Same clinical dataset and the same MobileNet backbone as the earlier study1, with broader augmentation and class weighting.

Average IoU
86.7%vs 84% published+2.7 pts
Weighted IoU
87.3%vs 85% published+2.3 pts
Pixel accuracy
92.9%vs 93% publishedon par

1 G. Aykol-Şahin et al., “Efficiency of oral keratinized gingiva detection and measurement based on convolutional neural network,” Journal of Periodontology, 2024. doi:10.1002/JPER.24-0151

What took the most work

  • Rare classes. Teeth and background cover most pixels; there are only 180 crack masks among 28,904. Inverse-frequency class weights in the loss raised cavity IoU from 0.15 to 0.54 and crack IoU from 0.03 to 0.15 without hurting the large classes.
  • Masks on the GPU. COCO polygons are rasterized straight into preallocated GPU tensors, which cut peak host RAM by nearly 80% and let mask generation and augmentation run in parallel on an RTX 4060.
  • Model choice. I compared YOLOv11n with DeepLabV3 on ResNet50 and MobileNetV3 backbones for accuracy, size and speed. YOLO was quick to integrate but weak on small lesions, with an mAP@0.5 of 0.195 for caries.
  • Training time. Learning-rate sweeps, mixed precision and a one-cycle cosine schedule got the mobile model past 95% pixel accuracy in under three hours on a consumer GPU.

On the phone

The model runs through TorchScript and PyTorch Lite with Vulkan. I tried quantization-aware training and dynamic int8 quantization, but the bigger gains came from the camera pipeline: YUV conversion, aspect-aware mask scaling and multithreaded overlay rendering. After that, almost all of a frame is the model itself, so the app lets you trade resolution for speed. Inference runs on its own thread next to the camera, so the preview never waits for the model.

Time per frame on a Galaxy S20 FE
1080p 518 ms
480 × 640 312 ms
  • Frame conversion 12 ms
  • Inference 488 ms
  • Mask overlay 16 ms

These are timings inside the running app, not isolated inference. The camera and the model run on separate threads and share the GPU, so the preview stays at 60 fps while the mask refreshes about every 300 ms.

DentAI photo mode: a molar with a large cavity, and the same photo with teeth, cavity and caries masks.
Photo mode
DentAI live mode: teeth highlighted over the camera preview.
Live camera
Per-class results and limitations
Mobile model, IoU per class
Teeth 95.66%
Background 83.97%
Cavities 54.04%
Caries 38.51%
Cracks 14.77%

Lesion scores are low for three reasons: very few examples, very small regions, and annotations that often mark only part of a lesion. In several images the prediction covers more of the lesion than the ground truth does, and pixel IoU counts that against it. Merged into one “abnormality” class, caries, cavities and cracks reach about 85% IoU.

The gingiva model reaches 86.95% IoU on teeth, 86.24% on keratinized and 86.98% on non-keratinized gingiva.

It’s a screening aid, not a diagnostic tool.

Projects

TEKNOFEST

2nd place, International UAV Competition

I was chief software engineer for our team’s drone, responsible for the vision and flight software. We placed second. Afterwards I mentored newer members of the GTU aviation club.

OpenCV, DroneKit, MAVLink, Gazebo, QGroundControl

Verify the certificate

Internship project

AirportGPT

An airport assistant I built at iGA Istanbul Airport. It answers from live flight data and a retrieval index of airport FAQs instead of model memory, and runs on a local LLM so nothing leaves the network.

FastAPI, LangChain, Next.js, RAG

Demo in the repository

Personal project

XGPrice

Forecasting trend changes in stock prices with XGBoost. I built the feature engineering, future-return targets, walk-forward evaluation and a backtesting engine with position sizing, stop-loss and take-profit rules. The backtests don’t model transaction costs, so I read them as experiments, not a track record.

Python, XGBoost, pandas, Optuna

Graduation project

360Vision

Lane detection and lateral distance estimation from a single camera. YOLOP segments the lane lines; the pipeline cleans the masks, fits polynomial lane boundaries and estimates the distance to each side with a camera model.

PyTorch, YOLOP, OpenCV

Developer tool

Microservice Orchestrator

A local dashboard for working across 20+ Spring Boot and React repositories. It switches branches without losing local config, gives each Git worktree its own development database, and streams service terminals to the browser.

Next.js, TypeScript, Docker, PostgreSQL

Experience

  1. Jul 2025 – now

    TÜBİTAK BİLGEM

    Software Engineer / Researcher

    Backend services for the national public financial management system: Java 21, Spring Boot, Kafka, PostgreSQL and React. Features across services, production debugging, code review and releases, plus release and review tooling now used by several teams.

  2. Jan 2025 – now

    Scale AI

    Contributor

    Reviewing and rating code written by frontier LLMs for RLHF training.

  3. Jan – Feb 2025

    iGA Istanbul Airport

    Intern

    Built AirportGPT.

  4. Oct 2023 – Jan 2024

    TÜBİTAK BİLGEM

    Part-time undergraduate researcher

    Dynamic logging with Spring AOP for a large online-operations product.

  5. Jul – Aug 2023

    TÜBİTAK BİLGEM

    Intern

    Personnel management system rewrite in Spring Boot and Next.js.

Education and languages

Gebze Technical University

B.Sc. Computer Engineering, 2021–2025
GPA 3.03 / 4.00

Gdańsk University of Technology

Erasmus+ exchange, Data Engineering, 2023

English

C2 · TOEFL iBT 118 / 120 (2026)

McKinsey Forward

Leadership and communication, 2024

Tools

Backend
Java 21, Spring Boot, Spring Data JPA, Hibernate, Kafka, PostgreSQL, Liquibase, OAuth2
Testing and delivery
JUnit 5, Mockito, Spock, Spring Cloud Contract, Jest, Jenkins, Maven, Docker, Kubernetes manifests, Nginx
ML and vision
PyTorch, DeepLabV3, YOLO, OpenCV, TorchScript, PyTorch Lite, CameraX, XGBoost, scikit-learn
LLM applications
FastAPI, LangChain, RAG, local inference, agent tooling
Frontend
React, Next.js, TypeScript
Languages
Java, Python, TypeScript, SQL, C, C++

Contact

Email is the fastest way to reach me.

salihkaragollu@gmail.com

Salih Karagöllü

This site counts link and button clicks as daily totals. No cookies, no stored IP addresses.

Java, Java 21, Spring Boot, Spring Framework, Python, JavaScript, TypeScript, React, Next.js, PostgreSQL, SQL, Kafka, Docker, Jenkins, ArgoCD, Pinpoint, Git, Bitbucket, GitHub, Microservices, REST API, OAuth2, CI/CD, DevOps, TensorFlow, PyTorch, Deep Learning, Machine Learning, Computer Vision, NLP, LLM, RAG, LangChain, YOLO, DeepLabV3, MobileNet, Quantization, Segmentation, CNNs, RLHF, XGBoost, Scikit-Learn, Pandas, Matplotlib, Data Augmentation, Data Preprocessing, Optuna, Jupyter, COCO, Object Detection, Image Processing, C, C++, HTML, CSS, Material UI, Shadcn, ASP.NET Core, Agile, Jira, Scrum, OOP, Spring Security, Hibernate, JPA, Kubernetes, AWS, EC2, Linux, Bash