TXT Group is an international, IT end-to-end provider of consultancy, software solutions and services, supporting the digital transformation of customers’ products and core processes. With a proprietary software portfolio and deep expertise in vertical domains, TXT Group operates across different markets, with a growing footprint in Aerospace, Aviation, Defense, Industrial, Government and Fintech. The holding company, TXT e-Solutions, has been listed on the Italian Stock Exchange - STAR segment (TXT.MI) - since July 2000. TXT Group is headquartered in Milan and has subsidiaries in Italy, Germany, the United Kingdom, France, Switzerland and the United States of America.
MLOps & Platform Engineer:
TXT Group, a company within the TXT Group, is seeking an MLOps & Platform Engineer to join its Industrial Business Unit.
The ideal candidate will be responsible to design, build and operate the company’s application and AI platform, ensuring secure, scalable and highly available environments for both enterprise applications and AI/ML workloads.
At least three years’ relevant experience in platform and infrastructure engineering is required, with production ownership of: Kubernetes, CI/CD pipelines and cloud infrastructure.
Main responsibilities:
Design and manage cloud-native platforms, Kubernetes clusters and containerised applications;
Build and maintain CI/CD pipelines, and automate infrastructure provisioning and application deployment;
Design, deploy and manage infrastructure for AI systems: LLM and embedding model serving, vector databases, application databases and caching systems;
Define and maintain CI/CD pipelines for AI applications and data pipelines, with reproducible environments and secure release strategies;
Manage versioning of models, prompts and configurations, and support fine-tuning and retraining pipelines where required;
Collaborate with software developers and AI engineers to streamline delivery and operations;
Implement monitoring, logging, tracing and alerting for LLM applications and data pipelines, covering latency, error rate, cost per call, drift and production quality metrics;
Ensure scalability, high availability and operational continuity of AI services, including knowledge base ingestion and update pipelines;
Optimise inference and embedding costs through resource sizing, quantisation, batching and infrastructure-level caching;
Ensure overall platform security, observability, performance and operational reliability;
Manage secrets, API keys, access control and environment isolation for AI services;
Support offline evaluation and A/B testing activities by providing the necessary infrastructure and telemetry data.
Indispensable technical skills:
Solid experience with containerisation and orchestration technologies such as Docker, Kubernetes, Helm and Docker Compose;
Proficiency with CI/CD and DevOps tooling, including Git, GitLab and GitLab CI/CD (or equivalents such as GitHub Actions), GitOps workflows, container registries and release management practices, with automation skills in Python and Bash;
Strong background in infrastructure automation on Linux, using Terraform and Ansible to implement Infrastructure as Code and manage virtualised environments;
Solid understanding of networking and security fundamentals (TCP/IP, HTTP/HTTPS, DNS), reverse proxies and ingress controllers (Nginx, Traefik), TLS/SSL, identity and access management (Keycloak, OAuth2/OpenID Connect), secrets management and IAM;
Experience with observability stacks such as Prometheus, Grafana, Loki, OpenTelemetry, OpenSearch and Alertmanager (or an equivalent ELK-based stack), applied to production ML and LLM systems;
Hands-on experience with AI/MLOps tooling, including MLflow and experiment tracking platforms such as Weights & Biases, model registries, and model serving frameworks such as vLLM, Ollama, Triton Inference Server, SageMaker or Vertex AI, applied to GPU-based inference workloads;
Experience deploying and operating vector databases (Qdrant, pgvector), embedding models and RAG pipelines as part of production AI inference services;
Proficiency in backend and data technologies, including Python, FastAPI and REST APIs, together with operational experience running PostgreSQL, Microsoft SQL Server and Redis in production;
Practical experience operating production databases across SQL/NoSQL and vector stores, covering provisioning, backup and scaling;
Familiarity with LLMOps practices, including prompt and experiment tracking, inference cost monitoring and model version management;
Understanding of inference optimisation techniques such as quantisation, batching, caching and GPU-level optimisation (e.g. TensorRT);
Experience managing the ML/LLM model lifecycle, including fine-tuning and retraining pipelines, offline evaluation and A/B testing;
Working knowledge of at least one major cloud platform (Microsoft Azure, AWS or Google Cloud Platform) and its services for ML/AI workloads.
Familiarity with HashiCorp Nomad is a PLUS.
Education: Bachelor’s or Master’s degree in Computer Science, Computer Engineering or a similar field.
The perfect candidate will also possess problem-solving skills, curiosity, the ability to work independently, a proactive approach, a sense of responsibility, a collaborative spirit, the ability to draft technical documentation, and the ability to work effectively within cross-functional teams.
Why choose TXT Group:
Hybrid working mode;
Career opportunities in a fast-growing, international and dynamic environment;
Continuous training on technical and project-related topics;
Corporate benefits including health insurance, welfare services, meal vouchers, and employee discounts;
The contract offered for this position is a permanent one, and the salary range for this position is between EUR 35.000 € and EUR 40.000 € gross per year.
The grading/level will be defined during the selection process based on the candidate's profile and in accordance with the National Collective Labor Agreement (Metalworking Industry).
Position open to candidates without distinction of gender, pursuant to Legislative Decree 198/2006.
The company promotes equal opportunities and values diversity in all its forms.
#LI-Hybrid