МенюКОМЭКСПО WORK
КОМЭКСПО
Возможности платформы

Курсы и специалистов можно выбирать без входа.

Ещё в КОМЭКСПО

Другие возможности

Нужна помощь?Напишите нам — поможем разобраться
КОМЭКСПО СкиллыОткрыть ассистента ↗
← Все навыки

ml pipeline workflow

Complete end-to-end MLOps pipeline orchestration from data preparation through model deployment.

Каталог из GitHub · что мы проверяем

Собираем навыки из Agentic Awesome Skills на GitHub. Это работы авторов сообщества, а не собственные разработки КОМЭКСПО. В карточках сохраняем источник и фиксированную версию.

Автоматические проверки

  • При импорте: проверяем адреса источников и убираем повторяющиеся идентификаторы. Некоторые категории риска исключаем.
  • При загрузке инструкции: проверяем формат SKILL.md, кодировку и объём. Ищем упоминания дополнительных файлов и инструментов.
  • Сверяем обозначение лицензии с разрешённым списком и учитываем известные исключения источника. Непонятные условия требуют отдельной проверки.

«Требует проверки» означает, что инструкция ещё не загружена и не проверена. Статусы совместимости не подтверждают безопасность, качество ответа или работу во всех моделях. Полный аудит кода, прав и тестирование каждого навыка не проводились. Скрипты не запускаем.

Авторы и условия использования

Оригинальные тексты коллекции заявлены под CC BY 4.0; у сторонних материалов могут быть другие условия. Открытый GitHub не означает отсутствие авторских прав. Сохраняйте авторство, ссылку на лицензию и отметки об изменениях.

Русские пояснения и промпт-обёртки подготовлены КОМЭКСПО. Оригинальные инструкции не переведены. Мы не связаны с GitHub или авторами навыков и не заявляем об их одобрении сервиса. Правообладателям: контакты — укажите карточку, оригинал и суть обращения.

Использовать навык

Без регистрации

Вставьте в свой ИИ-чат и замените последнюю строку своей задачей. Это инструкция, а не подключение новых инструментов.

Скачать SKILL.md ↓

Источник: sickn33/agentic-awesome-skills · CC-BY-4.0 · Оригинал ↗

Атрибуция для копирования и распространения

Если распространяете скачанный файл, приложите атрибуцию и требуемые лицензией уведомления автора. При изменении инструкции укажите свои изменения. Условия лицензии

Как это работает

Навык задаёт подход к задаче: например, как редактировать текст или проверять код. Вы выбираете его в чате — ИИ получает эту инструкцию вместе с вашим запросом.

Поддерживаются текст, код и создание сайтов. Навык не подключает новые модели, терминал, MCP или аккаунты. Для изображений и видео используются отдельные инструменты.

Совместимость
Нужны доп. инструменты
Версия
465ad05638fb · фиксируется при добавлении
Лицензия
CC-BY-4.0
Звёзды репозитория
47 140

Ограничения оригинала

  • В оригинале есть ссылки на файлы или внешние инструменты. Они не устанавливаются; в чате применяются только инструкции SKILL.md.
Оригинальная инструкция SKILL.md · 7 638 символов
---
name: ml-pipeline-workflow
description: "Complete end-to-end MLOps pipeline orchestration from data preparation through model deployment."
risk: critical
source: community
date_added: "2026-02-27"
---

# ML Pipeline Workflow

Complete end-to-end MLOps pipeline orchestration from data preparation through model deployment.

## Do not use this skill when

- The task is unrelated to ml pipeline workflow
- You need a different domain or tool outside this scope

## Instructions

- Clarify goals, constraints, and required inputs.
- Apply relevant best practices and validate outcomes.
- Provide actionable steps and verification.
- If detailed examples are required, open `resources/implementation-playbook.md`.

## Overview

This skill provides comprehensive guidance for building production ML pipelines that handle the full lifecycle: data ingestion → preparation → training → validation → deployment → monitoring.

## Use this skill when

- Building new ML pipelines from scratch
- Designing workflow orchestration for ML systems
- Implementing data → model → deployment automation
- Setting up reproducible training workflows
- Creating DAG-based ML orchestration
- Integrating ML components into production systems

## What This Skill Provides

### Core Capabilities

1. **Pipeline Architecture**
   - End-to-end workflow design
   - DAG orchestration patterns (Airflow, Dagster, Kubeflow)
   - Component dependencies and data flow
   - Error handling and retry strategies

2. **Data Preparation**
   - Data validation and quality checks
   - Feature engineering pipelines
   - Data versioning and lineage
   - Train/validation/test splitting strategies

3. **Model Training**
   - Training job orchestration
   - Hyperparameter management
   - Experiment tracking integration
   - Distributed training patterns

4. **Model Validation**
   - Validation frameworks and metrics
   - A/B testing infrastructure
   - Performance regression detection
   - Model comparison workflows

5. **Deployment Automation**
   - Model serving patterns
   - Canary deployments
   - Blue-green deployment strategies
   - Rollback mechanisms

### Reference Documentation

See the `references/` directory for detailed guides:
- **data-preparation.md** - Data cleaning, validation, and feature engineering
- **model-training.md** - Training workflows and best practices
- **model-validation.md** - Validation strategies and metrics
- **model-deployment.md** - Deployment patterns and serving architectures

### Assets and Templates

The `assets/` directory contains:
- **pipeline-dag.yaml.template** - DAG template for workflow orchestration
- **training-config.yaml** - Training configuration template
- **validation-checklist.md** - Pre-deployment validation checklist

## Usage Patterns

### Basic Pipeline Setup

```python
# 1. Define pipeline stages
stages = [
    "data_ingestion",
    "data_validation",
    "feature_engineering",
    "model_training",
    "model_validation",
    "model_deployment"
]

# 2. Configure dependencies
# See assets/pipeline-dag.yaml.template for full example
```

### Production Workflow

1. **Data Preparation Phase**
   - Ingest raw data from sources
   - Run data quality checks
   - Apply feature transformations
   - Version processed datasets

2. **Training Phase**
   - Load versioned training data
   - Execute training jobs
   - Track experiments and metrics
   - Save trained models

3. **Validation Phase**
   - Run validation test suite
   - Compare against baseline
   - Generate performance reports
   - Approve for deployment

4. **Deployment Phase**
   - Package model artifacts
   - Deploy to serving infrastructure
   - Configure monitoring
   - Validate production traffic

## Best Practices

### Pipeline Design

- **Modularity**: Each stage should be independently testable
- **Idempotency**: Re-running stages should be safe
- **Observability**: Log metrics at every stage
- **Versioning**: Track data, code, and model versions
- **Failure Handling**: Implement retry logic and alerting

### Data Management

- Use data validation libraries (Great Expectations, TFX)
- Version datasets with DVC or similar tools
- Document feature engineering transformations
- Maintain data lineage tracking

### Model Operations

- Separate training and serving infrastructure
- Use model registries (MLflow, Weights & Biases)
- Implement gradual rollouts for new models
- Monitor model performance drift
- Maintain rollback capabilities

### Deployment Strategies

- Start with shadow deployments
- Use canary releases for validation
- Implement A/B testing infrastructure
- Set up automated rollback triggers
- Monitor latency and throughput

## Integration Points

### Orchestration Tools

- **Apache Airflow**: DAG-based workflow orchestration
- **Dagster**: Asset-based pipeline orchestration
- **Kubeflow Pipelines**: Kubernetes-native ML workflows
- **Prefect**: Modern dataflow automation

### Experiment Tracking

- MLflow for experiment tracking and model registry
- Weights & Biases for visualization and collaboration
- TensorBoard for training metrics

### Deployment Platforms

- AWS SageMaker for managed ML infrastructure
- Google Vertex AI for GCP deployments
- Azure ML for Azure cloud
- Kubernetes + KServe for cloud-agnostic serving

## Progressive Disclosure

Start with the basics and gradually add complexity:

1. **Level 1**: Simple linear pipeline (data → train → deploy)
2. **Level 2**: Add validation and monitoring stages
3. **Level 3**: Implement hyperparameter tuning
4. **Level 4**: Add A/B testing and gradual rollouts
5. **Level 5**: Multi-model pipelines with ensemble strategies

## Common Patterns

### Batch Training Pipeline

```yaml
# See assets/pipeline-dag.yaml.template
stages:
  - name: data_preparation
    dependencies: []
  - name: model_training
    dependencies: [data_preparation]
  - name: model_evaluation
    dependencies: [model_training]
  - name: model_deployment
    dependencies: [model_evaluation]
```

### Real-time Feature Pipeline

```python
# Stream processing for real-time features
# Combined with batch training
# See references/data-preparation.md
```

### Continuous Training

```python
# Automated retraining on schedule
# Triggered by data drift detection
# See references/model-training.md
```

## Troubleshooting

### Common Issues

- **Pipeline failures**: Check dependencies and data availability
- **Training instability**: Review hyperparameters and data quality
- **Deployment issues**: Validate model artifacts and serving config
- **Performance degradation**: Monitor data drift and model metrics

### Debugging Steps

1. Check pipeline logs for each stage
2. Validate input/output data at boundaries
3. Test components in isolation
4. Review experiment tracking metrics
5. Inspect model artifacts and metadata

## Next Steps

After setting up your pipeline:

1. Explore **hyperparameter-tuning** skill for optimization
2. Learn **experiment-tracking-setup** for MLflow/W&B
3. Review **model-deployment-patterns** for serving strategies
4. Implement monitoring with observability tools

## Related Skills

- **experiment-tracking-setup**: MLflow and Weights & Biases integration
- **hyperparameter-tuning**: Automated hyperparameter optimization
- **model-deployment-patterns**: Advanced deployment strategies

## Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
✦ Чат AI