New RunLLM response generated: >>>*# Question* An...
# runllm-flyte-trial
a
New RunLLM response generated: >>>*# Question* Another question about scale. For a test we have ~70K workflows running. We've scaled up the propeller workers and other settings - see
yml
at the end of this post. What I'm observing seems a bit strange. When flyte first starts up it starts to schedule a lot of pods. I see up to 10K as
pending
but once the
unprocessed queue depth
goes down the system is scheduling hardly any work. I would expect it to try to schedule 10s of thousands of pods with 70K workflows running. The roundtrip latency looks ok to my untrained eyes. We checked our k8s API and can't find any throttling going on there - but maybe we are looking at the wrong things. On the other hand I also see some tasks that are marked as
Running
for over 1 hour with the status message
Sent to K8s...
but no pod log link is available yet (see image below as well)
Copy code
inline:
      catalogcache:
        reader:
          maxItems: 100000
          maxRetries: 3
          workers: 100
        writer:
          maxItems: 100000
          maxRetries: 3
          workers: 100
      workqueue:
        workers: 1200
        maxItems: 100000
        config:
          workers: 1200
          maxItems: 100000
      workqueue.Config:
        workers: 1200
        maxItems: 100000
      workqueueConfig:
        workers: 1200
        maxItems: 100000
    # Expose the metrics port for the Prometheus Service Monitor
    propeller:
      workers: 1800
      prof-port: 10254
      metrics-prefix: "flyte:"
      max-workflow-retries: 50
      kube-client-config:
        qps: 4000 # Refers to max rate of requests (queries per second) to kube-apiserver
        burst: 1200 # refers to max burst rate.
        timeout: 30s # Refers to timeout when talking with the kube-apiserver
      max-streak-lenght: 2
      event:
        rate: 10000
        capacity: 20000
        max-retries: 10
>>
# Answer It seems like you might be encountering issues related to Kubernetes API server throttling or FlytePropeller configuration. According to the Flyte performance documentation, even with increased
qps
and
burst
values, the Kube API server can still throttle requests transparently, leading to performance drops. Monitoring specific metrics like
workqueue_depth
can help identify if throttling is the issue. Additionally, the FlytePropeller configuration suggests that increasing the number of workers and adjusting the
qps
and
burst
values can help, but it requires careful tuning. The
round_latency
metric is crucial for understanding the performance of FlytePropeller. Would you like more detailed steps on how to monitor and adjust these settings? Sources: - Flyte performance documentation - FlytePropeller configuration Slack Link https://flyte-org.slack.com/archives/CP2HDHKE1/p1725456376.128309 👍 0 button 👎 0 button 💡 Hint: Mention @runllm in the thread for followups.