<?xml version="1.0" encoding="UTF-8"?>
<feed xml:lang="en-US" xmlns="http://www.w3.org/2005/Atom">
  <id>tag:status.starsling.dev,2005:/history</id>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev"/>
  <link rel="self" type="application/atom+xml" href="https://status.starsling.dev/history.atom"/>
  <title>starslingdev Status - Incident history</title>
  <updated>2026-09-01T15:00:22.096+00:00</updated>
  <author>
    <name>starslingdev</name>
  </author>
  
<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmtisqvk8a4re0kpw4ihff1ok</id>
  <published>2026-09-01T15:00:22.096+00:00</published>
  <updated>2026-09-01T15:00:46.084+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmtisqvk8a4re0kpw4ihff1ok"/>
  <title>Delays in commit processing</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 2 days, 12 hours and 59 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Pull Requests</p>
    <p><small>Sep <var data-var='date'> 1</var>, <var data-var='time'>15:00:46</var> GMT+0</small><br /><strong>Investigating</strong> -
  Diffs in the PR view may be stale for several minutes. We are investigating and scaling up resources..</p>
<p><small>Sep <var data-var='date'> 1</var>, <var data-var='time'>15:00:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Pull Requests.</p>
<p><small>Sep <var data-var='date'> 1</var>, <var data-var='time'>16:00:50</var> GMT+0</small><br /><strong>Investigating</strong> -
  Time to update pull request diffs have improved to normal thresholds..</p>
<p><small>Sep <var data-var='date'> 1</var>, <var data-var='time'>16:01:21</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmtap4t4x016h0qqmcfs0uysw</id>
  <published>2026-08-26T22:56:31.793+00:00</published>
  <updated>2026-08-26T22:57:08.551+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmtap4t4x016h0qqmcfs0uysw"/>
  <title>Incident with Actions and Pull Requests</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 days, 18 hours and 12 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Pull Requests, Third Party: GitHub → Actions</p>
    <p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>22:57:08</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating elevated delays and timeouts affecting Actions workflow runs triggered by pull request events. 20% of actions runs have delayed starts of more than 5 minutes and up to 4% of runs failed to trigger. We are actively working on mitigation and will provide updates as we learn more..</p>
<p><small>Aug <var data-var='date'> 27</var>, <var data-var='time'>00:01:05</var> GMT+0</small><br /><strong>Investigating</strong> -
  We&#039;ve applied mitigations and are seeing recovery in Actions workflow runs and blocked pull request merges. We&#039;re continuing to monitor for sustained health of merge commit creates before resolving..</p>
<p><small>Aug <var data-var='date'> 27</var>, <var data-var='time'>00:26:05</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>22:56:31</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions and Pull Requests.</p>
<p><small>Aug <var data-var='date'> 27</var>, <var data-var='time'>00:25:58</var> GMT+0</small><br /><strong>Investigating</strong> -
  We confirmed full recovery beginning at 23:58 UTC. Actions workflow runs and pull request merges are operating normally. We will now resolve the incident while continuing to monitor service health..</p>
<p><small>Aug <var data-var='date'> 27</var>, <var data-var='time'>00:26:44</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed. &lt;br /&gt;&lt;br /&gt;Background jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information. &lt;br /&gt;&lt;br /&gt;We reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations. &lt;br /&gt;&lt;br /&gt;We are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmta8j4dgipom0kmgshhe1emw</id>
  <published>2026-08-26T15:11:58.248+00:00</published>
  <updated>2026-08-26T15:12:21.306+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmta8j4dgipom0kmgshhe1emw"/>
  <title>Incident with Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 7 days, 1 hour and 32 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>15:12:21</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>15:11:58</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded availability for Actions.</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>15:23:10</var> GMT+0</small><br /><strong>Investigating</strong> -
  We&#039;ve identified an issue with a database primary and are failing over to a replica immediately.</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>15:48:07</var> GMT+0</small><br /><strong>Investigating</strong> -
  primary failover briefly improved performance but did not fully mitigate, we&#039;ve throttled inbound traffic and are investigating upstream Vitess issues.</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>18:00:48</var> GMT+0</small><br /><strong>Monitoring</strong> -
  All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>18:01:30</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system.  &lt;br /&gt;&lt;br /&gt;At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. &lt;br /&gt;&lt;br /&gt;3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. &lt;br /&gt;&lt;br /&gt;Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27. &lt;br /&gt;&lt;br /&gt;Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs. &lt;br /&gt;&lt;br /&gt;Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>16:14:16</var> GMT+0</small><br /><strong>Investigating</strong> -
  We believe we&#039;ve identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn&#039;t recur. Some customers will continue to see delays as we ramp up..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>16:49:07</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is operating normally..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>16:50:28</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>17:32:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to observe recovery and expect actions inbound queues to be back to normal in &lt;30min. Work will continue to flow through the system subject to per-customer concurrency limits..</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>17:54:33</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt7ayowafyaw0rmgdlehel0y</id>
  <published>2026-08-24T13:56:54.964+00:00</published>
  <updated>2026-08-24T13:56:55.051+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt7ayowafyaw0rmgdlehel0y"/>
  <title>Actions delays in starting runs</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 13 hours and 47 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>13:56:55</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>14:22:59</var> GMT+0</small><br /><strong>Investigating</strong> -
  Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery..</p>
<p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>14:26:11</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>14:34:42</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 24, 2026, between 13:33 UTC and 14:04 UTC,  3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. &lt;br /&gt; &lt;br /&gt;The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. &lt;br /&gt;&lt;br /&gt;To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szr9n2o0q0rmgg6lnp1fl</id>
  <published>2026-08-17T13:40:03.620+00:00</published>
  <updated>2026-08-17T14:04:45.634+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szr9n2o0q0rmgg6lnp1fl"/>
  <title>Incident with GitHub.com</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 18 days, 23 hours and 43 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests, Third Party: GitHub → Webhooks, Third Party: GitHub → Pull Requests, Third Party: GitHub → Actions</p>
    <p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:04:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:40:03</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of impacted performance for some GitHub services..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:41:47</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:42:42</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:44:02</var> GMT+0</small><br /><strong>Investigating</strong> -
  Webhooks is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:45:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available.</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>15:01:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:46:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:58:14</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:24:53</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:31:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  Copilot is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:45:53</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>21:15:46</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. &lt;br /&gt;&lt;br /&gt;Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. &lt;br /&gt;&lt;br /&gt;The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. &lt;br /&gt;&lt;br /&gt;The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. &lt;br /&gt;&lt;br /&gt;Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery. &lt;br /&gt;&lt;br /&gt;Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints. &lt;br /&gt;&lt;br /&gt;To prevent recurrence, our follow-up actions include: &lt;br /&gt;&lt;br /&gt;- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity. &lt;br /&gt;&lt;br /&gt;- Auditing Istio request, concurrency, and scaling limits across affected services. &lt;br /&gt;&lt;br /&gt;- Reviewing retry limits and backoff behavior across gateways and clients. &lt;br /&gt;&lt;br /&gt;- Addressing the VS Code retry behavior that amplified Copilot token traffic. &lt;br /&gt;&lt;br /&gt;- Improving load-balancer capacity monitoring and regional failover safeguards..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:49:24</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:54:14</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:58:09</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:58:12</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:58:17</var> GMT+0</small><br /><strong>Investigating</strong> -
  Webhooks is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>15:10:00</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>15:21:36</var> GMT+0</small><br /><strong>Investigating</strong> -
  Git Operations is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>15:40:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  Webhooks is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>15:42:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>16:16:13</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>16:36:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>16:59:38</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>17:34:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>17:36:29</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>18:11:30</var> GMT+0</small><br /><strong>Investigating</strong> -
  We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>18:23:47</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>18:48:19</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>19:01:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is operating normally..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>20:22:28</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is operating normally..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>19:13:20</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>20:08:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>20:45:28</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>17:30:49</var> GMT+0</small><br /><strong>Investigating</strong> -
  Git Operations is experiencing degraded performance. We are continuing to investigate..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szr9k2o0n0rmg69cp9hs3</id>
  <published>2026-08-13T14:45:40.422+00:00</published>
  <updated>2026-08-13T14:45:40.503+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szr9k2o0n0rmg69cp9hs3"/>
  <title>Incident with Webhooks</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 2 days, 2 hours and 54 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Webhooks, Third Party: GitHub → Pull Requests</p>
    <p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:45:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Webhooks.</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:46:11</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:46:23</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:46:52</var> GMT+0</small><br /><strong>Investigating</strong> -
  Git Operations is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:56:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  Packages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>14:58:49</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are currently investigating a brief degradation of service for Git operations (specifically pushes), issues, pull requests, package registry, and webhooks between 14:32 and 14:46 UTC. We have identified the source of the degradation and are investigating mitigation strategies to prevent recurrence..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>15:33:02</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Git Operations, Issues, Packages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>15:33:54</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We have temporarily disabled a background job which caused the impact. At this time the impact is fully mitigated..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>15:36:34</var> GMT+0</small><br /><strong>Resolved</strong> -
  Between 14:24 and 14:53 UTC on 13 August 2026, a routine background job to delete an organization overwhelmed a key shared database, causing multiple GitHub services to briefly return elevated errors and slower responses. Most affected was the webhook management API, with smaller impact to Git operations, pull requests, issues, packages, sign-in, and Copilot. Impact cleared on its own at about 14:53 UTC once the job finished; we resolved the incident at 15:36 UTC. &lt;br /&gt;&lt;br /&gt;Affected users may have experienced a brief increase in errors and slower responses, primarily when creating, listing, or updating webhooks, with smaller impacts to pull requests, issues, packages, and Git operations. Failures peaked at about 1% for several minutes around 14:37 UTC. &lt;br /&gt;&lt;br /&gt;To prevent future incidents, we&#039;ve already shipped an update that turns on the safer deletion path for organizations, along with caps on deletion holds on databases. Building on these changes, we&#039;re auditing all bulk deletion and cleanup jobs that write to shared databases to prevent similar issues in future..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szr9l2o0p0rmgen9imoyl</id>
  <published>2026-08-12T16:16:17.994+00:00</published>
  <updated>2026-08-12T16:16:18.105+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szr9l2o0p0rmgen9imoyl"/>
  <title>Incident with Pull Requests and Issues</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day and 1 hour</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Pull Requests</p>
    <p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:16:18</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Issues and Pull Requests.</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:24:30</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of errors affecting Pull Requests and Issues on GitHub.com. Some users may encounter 500 errors when loading pull request and issue pages. Our engineering teams are actively investigating the root cause, which appears to be related to a database infrastructure issue. We will provide an update as soon as we have more information..</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:35:40</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Issues and Pull Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:38:38</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We identified the source of errors affecting Pull Requests, Issues, and Search on GitHub.com and have applied a mitigation. A database index hint was referencing an index that had been removed by a recent migration, causing query failures for some users. We disabled the problematic configuration and are seeing recovery across affected services. We are continuing to monitor to confirm full resolution..</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:41:18</var> GMT+0</small><br /><strong>Resolved</strong> -
  Between 16:03 and 16:29 UTC on August 12, some users encountered errors when viewing pull requests, issues, and search results. During this period, about 1.9% of Pull Request requests and 0.9% of Issues requests failed. During a database migration, two indexes were removed while application settings still referenced them, causing affected requests to fail. We detected the issue after the migration reached one database shard and before it progressed to the remaining shards. We restored service by disabling both settings. We are improving safeguards around database migrations and application configuration to prevent similar mismatches from causing errors..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szr922o0j0rmgxvj2tlp3</id>
  <published>2026-08-11T14:50:40.219+00:00</published>
  <updated>2026-08-11T14:50:40.309+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szr922o0j0rmgxvj2tlp3"/>
  <title>Incident with GraphQL API Requests</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 13 days, 4 hours and 16 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests</p>
    <p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>14:50:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for API Requests.</p>
<p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>15:28:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of a small increase in error rates affecting GraphQL API requests. We are working on increasing capacity and continue to investigate the increased errors. We will provide another update when we have more information.</p>
<p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>16:49:24</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have returned to a healthy baseline on GraphQL API requests. We will continue to work on investigations into the errors seen during this incident..</p>
<p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>16:49:34</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>20:06:50</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We have identified and mitigated increased error rates affecting GraphQL API requests. A fix to increase service capacity has been deployed and error rates have returned to normal levels. We are resolving this incident..</p>
<p><small>Aug <var data-var='date'> 11</var>, <var data-var='time'>20:06:56</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 11, 2026, between 14:00 UTC and 16:00 UTC the GraphQL API service was degraded and customers in  saw higher than normal timeouts. On average, the timeout rate was 0.06% and peaked at 0.14% of requests routing to the service. &lt;br /&gt;&lt;br /&gt;This was due to increased utilization at one of our sites which caused resource contention across our dependencies, leading to an increase in timeouts for GraphQL requests. We mitigated the incident by increasing capacity to alleviate the capacity bottleneck.  &lt;br /&gt;&lt;br /&gt;We are working to improve our monitoring so that we can proactively reduce the impact of high consumption requests in addition to scaling up; Additionally, we will improve our time to detection and mitigation of issues like this one in the future..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szyyd2ogm0rmgvz7uvj8h</id>
  <published>2026-08-06T15:22:49.021+00:00</published>
  <updated>2026-08-07T00:01:26.921+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szyyd2ogm0rmgvz7uvj8h"/>
  <title>Incident with Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 26 days, 17 hours and 55 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>00:01:26</var> GMT+0</small><br /><strong>Investigating</strong> -
  System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.&lt;br /&gt;&lt;br /&gt;Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.&lt;br /&gt;&lt;br /&gt;We are monitoring all affected services for sustained recovery and will provide another update shortly..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>00:05:05</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>00:06:24</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>00:59:01</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We’re investigating reports that some Actions Runner Controller runners are taking longer than expected to recover. We’ll provide an update as our investigation progresses..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>02:04:44</var> GMT+0</small><br /><strong>Resolved</strong> -
  On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes. &lt;br /&gt;&lt;br /&gt;The incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services. &lt;br /&gt;&lt;br /&gt;These services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events. &lt;br /&gt;&lt;br /&gt;As the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the services responsible for job assignment, runners were getting assigned jobs that were no longer valid and then getting stuck retrying those jobs, preventing them from picking up valid work. &lt;br /&gt;&lt;br /&gt;This second stage of impact was mitigated by deploying changes to prevent runners from repeatedly attempting to acquire invalid jobs. These mitigations allowed the accumulated queues to drain and Actions to recover to normal operation. &lt;br /&gt;&lt;br /&gt;Some Actions Runner Controller (ARC) runners remained stuck after the incident. A mitigation deployed during the incident inadvertently affected these runners, causing some to remain offline until they were manually recovered. We subsequently rolled back the change and are adding automatic recovery in upcoming Runner and ARC releases. &lt;br /&gt;&lt;br /&gt;Some jobs created during the incident were also left stuck unable to be retried or canceled.  CLI and UI solutions for customers to address these were shared at https://github.com/orgs/community/discussions/204152#discussioncomment-17946043. &lt;br /&gt;&lt;br /&gt;To prevent recurrence, we are making improvements to deployment and capacity safeguards for the affected services, strengthening monitoring for the conditions that preceded the incident, improving the resiliency and recovery of queued work and runner assignment, and adding automatic recovery for self-hosted runners affected by similar failure conditions. We are also making additional improvements to reduce the risk of cascading failures and accelerate recovery during large-scale Actions disruptions..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>02:03:41</var> GMT+0</small><br /><strong>Monitoring</strong> -
  During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners.&lt;br /&gt;&lt;br /&gt;The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future.&lt;br /&gt;&lt;br /&gt;Some workflow-triggering events, including push and pull request events, were not processed during the incident and cannot be replayed automatically. Customers may need to repeat the triggering action by pushing a new commit, updating the pull request, or manually re-running the workflow where applicable..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>17:02:43</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on the issue affecting GitHub Actions. &lt;br /&gt;&lt;br /&gt;Workflow runs are still failing or delayed in starting, and some queued jobs may time out. &lt;br /&gt;&lt;br /&gt;Some requests to the Actions API are returning errors. Customers running migrations with GitHub Enterprise Importer may see failures. &lt;br /&gt;&lt;br /&gt;Our engineers have applied several mitigations and are rolling out a further fix now..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>21:30:42</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting GitHub Actions. Webhook triggers remain throttled to aid recovery, so many push and pull request events are not triggering new workflow runs.&lt;br /&gt;&lt;br /&gt;We identified runners being assigned jobs that are no longer valid and are deploying a change to address this issue. Both GitHub-hosted and self-hosted runners are affected.&lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, and GitHub Pages may experience failures or delays. Migrations using GitHub Enterprise Importer have been paused to support mitigation efforts..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>17:40:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting multiple GitHub services.&lt;br /&gt;&lt;br /&gt;Workflow runs are failing or delayed in starting, and some queued jobs may time out. &lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer might also affected. &lt;br /&gt;&lt;br /&gt;Webhook deliveries may be delayed. &lt;br /&gt;&lt;br /&gt;Engineers have applied a number of mitigations and are rolling out a further fix across all affected systems now..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>20:34:17</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.&lt;br /&gt;&lt;br /&gt;We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.&lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected..</p>
<p><small>Aug <var data-var='date'> 7</var>, <var data-var='time'>00:01:15</var> GMT+0</small><br /><strong>Investigating</strong> -
  &lt;br /&gt;System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.&lt;br /&gt;&lt;br /&gt;Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.&lt;br /&gt;&lt;br /&gt;We are monitoring all affected services for sustained recovery and will provide another update shortly..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>18:11:41</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting multiple GitHub services. &lt;br /&gt;&lt;br /&gt;Workflow runs are still failing or delayed in starting, and some queued jobs may time out. &lt;br /&gt;&lt;br /&gt;Customers using self-hosted runners may see errors or rate limiting when runners register. &lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer may also be affected. &lt;br /&gt;&lt;br /&gt;Webhook deliveries may be delayed.&lt;br /&gt;&lt;br /&gt;Engineers have applied further mitigations and are continuing to work towards full recovery..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>18:46:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting multiple GitHub services. &lt;br /&gt;&lt;br /&gt;Workflow runs are still failing, and jobs may remain queued for an extended period before starting or may time out. Jobs using GitHub-hosted runners are particularly affected while capacity is constrained. &lt;br /&gt;&lt;br /&gt;Customers using self-hosted runners may see errors or rate limiting when runners register. &lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. &lt;br /&gt;&lt;br /&gt;Recovery is taking longer than we expected, and engineers remain actively engaged..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>22:18:09</var> GMT+0</small><br /><strong>Investigating</strong> -
  We continue to make progress on the issue affecting GitHub Actions. We have deployed a fix that addresses runners being assigned jobs that are no longer valid, and are seeing improvement in job completion rates. For workflow runs that are starting, success rates have increased significantly and are now at 97%. Standard and larger runners are now draining queued work. A change is also in progress to mitigate issues with existing self-hosted runners that are not picking up jobs.&lt;br /&gt;&lt;br /&gt;Webhook triggers remain throttled to support recovery. Many push and pull request events are not yet triggering new workflow runs, and we are working to safely restore full throughput.&lt;br /&gt;&lt;br /&gt;GitHub Pages, Copilot code review, and Copilot coding agent may still experience failures or delays. Migrations using GitHub Enterprise Importer remain paused.&lt;br /&gt;&lt;br /&gt;We are continuing to monitor recovery and will provide another update as conditions improve..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>19:43:21</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on an issue affecting GitHub Actions. &lt;br /&gt;&lt;br /&gt;Capacity remains constrained and jobs may still be delayed or fail while it recovers gradually. Customers using self-hosted runners may see errors or rate limiting when runners register.  &lt;br /&gt;&lt;br /&gt;Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. &lt;br /&gt;&lt;br /&gt;Our engineers remain actively engaged..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>23:13:30</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have deployed fixes that address runners being assigned invalid jobs and are taking additional steps to clear the backlog of affected jobs. Job completion rates for running workflows have improved significantly, with success rates now at 99%. Global queues for hosted runner assignment are nearly burned down and concurrency queues for customers are being processed. Another change was deployed to accelerate processing the backlog of job requests.&lt;br /&gt;&lt;br /&gt;We are gradually restoring throughput for webhook-triggered Actions workflows and monitoring system stability. We have deployed a fix for self-hosted runners that were not picking up jobs and are enabling it incrementally.&lt;br /&gt;&lt;br /&gt;GitHub Pages, Copilot code review, and Copilot coding agent may still experience intermittent failures or delays. Migrations using GitHub Enterprise Importer remain paused.&lt;br /&gt;&lt;br /&gt;We continue to monitor recovery across all affected services and will provide another update as conditions improve..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>15:22:49</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>16:27:01</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>15:41:10</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>15:45:14</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating errors affecting GitHub Actions. Some workflow runs are failing to start or failing partway through, and some requests to the Actions REST API are returning errors. &lt;br /&gt;&lt;br /&gt;Some customers may also see unexpected rate limiting in their workflows. &lt;br /&gt;&lt;br /&gt;Engineers have identified the source of the disruption and are actively working on a mitigation.</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>15:53:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>16:19:30</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is operating normally..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>16:27:47</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on the issue affecting GitHub Actions. &lt;br /&gt;&lt;br /&gt;Some workflow runs are still delayed or failing to complete, and some requests to the Actions API are returning errors. &lt;br /&gt;&lt;br /&gt;Customers running migrations with GitHub Enterprise Importer may also see failures. &lt;br /&gt;&lt;br /&gt;Engineers are actively working towards full recovery..</p>
<p><small>Aug <var data-var='date'> 6</var>, <var data-var='time'>16:33:31</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions and Pages are experiencing degraded availability. We are continuing to investigate..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0szyyf2ogo0rmg86duag60</id>
  <published>2026-07-29T15:26:24.385+00:00</published>
  <updated>2026-07-29T15:26:24.457+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0szyyf2ogo0rmg86duag60"/>
  <title>Incident with Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 10 hours and 29 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 29</var>, <var data-var='time'>15:26:24</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded availability for Actions.</p>
<p><small>Jul <var data-var='date'> 29</var>, <var data-var='time'>15:34:07</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating an issue affecting GitHub Actions. Some customers may experience timeouts or failures with runner registration and workflow runs may be delayed during startup. Our team is actively working to mitigate the impact by scaling capacity across additional infrastructure..</p>
<p><small>Jul <var data-var='date'> 29</var>, <var data-var='time'>15:40:35</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 29</var>, <var data-var='time'>16:00:54</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 29, 2026, from 14:51 UTC to 15:28 UTC, GitHub Actions experienced elevated REST API request timeouts and errors, failures registering runners, and delayed workflow run starts for customers whose traffic was served by a single infrastructure site. This was caused by an under-provisioned internal Actions service in that site: under increased load its instances ran out of memory and became unresponsive, and because Actions API requests wait synchronously on that service, requests routed through the affected site stalled and timed out. During the incident, approximately 2% of workflows were delayed. Requests served by other sites remained unaffected. Both standard and larger hosted runners routed through the affected site could see delayed job starts. &lt;br /&gt;&lt;br /&gt;The issue was mitigated by scaling out the runner-administration service in the affected site and increasing the replica count, which restored API availability and returned workflow run starts to normal. We are working to add horizontal autoscaling, memory-saturation alerting, and scaling-forecast monitoring for this service, along with responder playbooks, to reduce the likelihood of similar issues in the future..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065k2oui0rmgskb880s0</id>
  <published>2026-07-27T03:53:19.028+00:00</published>
  <updated>2026-07-27T03:53:19.087+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065k2oui0rmgskb880s0"/>
  <title>Incident with GraphQL API Requests</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 15 hours and 51 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests</p>
    <p><small>Jul <var data-var='date'> 27</var>, <var data-var='time'>03:53:19</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for API Requests.</p>
<p><small>Jul <var data-var='date'> 27</var>, <var data-var='time'>04:09:10</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 26, 2026 at 21:34 UTC we began seeing intermittent errors on the GitHub GraphQL API. A subset of GraphQL API requests returned HTTP 502 errors in short bursts. During the impact window an average of 0.09% of GraphQL API requests in the affected region failed, with a peak of 0.50% of requests failing during the worst two-minute period at 03:02 UTC on July 27. Requests that failed generally succeeded when retried, and no data was lost or altered. Other GitHub services were not affected.&lt;br /&gt;&lt;br /&gt;The errors were traced to a single group of servers handling a share of GraphQL API traffic. Application processes on that group intermittently closed connections before completing responses. Impact ended at 03:52 UTC on July 27 when those processes were replaced, and we resolved the incident at 04:09 UTC on July 27 after confirming error rates had returned to normal.&lt;br /&gt;&lt;br /&gt;We are still investigating why those processes closed connections, and that work is being carried out by the team that owns the underlying compute platform. In the meantime we are adding detection and automated mitigation for when a single group of servers behaves differently from its peers..</p>
<p><small>Jul <var data-var='date'> 27</var>, <var data-var='time'>04:09:01</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065h2oue0rmgyvcxp6qa</id>
  <published>2026-07-25T12:31:50.953+00:00</published>
  <updated>2026-07-25T12:34:26.839+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065h2oue0rmgyvcxp6qa"/>
  <title>Actions run failures and delays</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 17 hours and 30 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>12:34:26</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are experiencing issues with GitHub Actions that are causing workflow run failures and delays for some users. Our engineering team is actively investigating the infrastructure issue and working to restore full functionality..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>12:58:32</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have applied a mitigation for the infrastructure issue affecting GitHub Actions. Workflow run failures and delays are improving but not yet fully resolved. Our engineering team continues to work on restoring full functionality across all affected infrastructure..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>12:59:58</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>13:12:41</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We have seen recovery in GitHub Actions performance following our earlier mitigation. Workflow runs are processing normally, though jobs queued before 12:40 UTC may still experience failures and will need to be retried..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>13:13:21</var> GMT+0</small><br /><strong>Resolved</strong> -
  Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/s65j9gslmfm8.</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>12:31:51</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded availability for Actions.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065o2oup0rmguoy79j0v</id>
  <published>2026-07-25T08:59:37.051+00:00</published>
  <updated>2026-07-25T08:59:37.122+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065o2oup0rmguoy79j0v"/>
  <title>Incident with Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 1 hour and 53 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>08:59:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>09:13:09</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>09:20:23</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We identified an issue causing delays in GitHub Actions run starts. Some users may have experienced longer than expected wait times when triggering workflow runs. We have applied mitigations and have recovered. Our team continues to monitor and investigate the root cause..</p>
<p><small>Jul <var data-var='date'> 25</var>, <var data-var='time'>09:25:30</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 25, 2026, GitHub Actions experienced two related periods of degradation that caused some workflow runs to be delayed by more than 5 minutes or end with infrastructure failures. &lt;br /&gt;&lt;br /&gt;First period (08:45 – 09:13 UTC): During planned maintenance on a critical-path Redis cluster for Actions, one participating region was left in a degraded state. Separately, an independent capacity operation temporarily removed another region from the cluster and redirected its traffic to the degraded region. This created cross-region inconsistencies in job-assignment state, causing workflow runs to be delayed, exhaust retries, or fail outright. At peak, about 7% of runs were delayed by more than 5 minutes, and 25% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 09:13 UTC by returning traffic to its normal distribution. &lt;br /&gt;&lt;br /&gt;Second period (12:08 – 12:48 UTC): As part of mitigating the first incident, traffic was returned to the regional instance that was still undergoing its capacity increase. Multiple Redis nodes in the scaling region experienced failures, increasing traffic to healthy nodes and causing connection limits to be reached on many nodes. At peak, 30% of runs were delayed by more than 5 minutes, and 60% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 12:48 UTC by redirecting workflow traffic away from the scaling region. &lt;br /&gt;&lt;br /&gt;We are adding stronger regional health and capacity checks before maintenance and requiring a stable observation period before restoring traffic. We are also improving automated connection resiliency, and partnering with our platform dependency to automatically detect and remediate unhealthy cluster members and shard imbalance.  More generally, we already had work underway to improve the resiliency and scale of this piece of Actions infrastructure..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065k2ouk0rmgmhlm5o1s</id>
  <published>2026-07-24T19:37:28.071+00:00</published>
  <updated>2026-07-24T19:37:28.159+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065k2ouk0rmgmhlm5o1s"/>
  <title>Incident with Pull Requests</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 21 hours and 42 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Pull Requests</p>
    <p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>19:37:28</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Pull Requests.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>19:40:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating errors creating pull requests.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>19:43:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>19:59:57</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have applied a mitigation and are monitoring for recovery.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>20:02:39</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>20:23:10</var> GMT+0</small><br /><strong>Resolved</strong> -
  Between July 24, 19:17 UTC and July 24, 20:02 UTC, users were unable to create pull requests due to a database schema change. In total, 113,930 pull request creation attempts were impacted across 50,904 users, with an average error rate of 1.75% and a maximum error rate of 2.25% for all requests to Pull Requests service. Existing pull requests and other GitHub functionality were not affected. The issue was resolved by reverting the change to the affected database, upon which pull request creation immediately resumed.&lt;br /&gt;&lt;br /&gt;The root cause was related to a backfill workflow into the Vitess keyspace hosting Pull Request data. The backfill Vitess command encountered errors and increased VReplication lag, and the workflow was canceled at 19:17 UTC. The cancellation executed a misunderstood Vitess codepath that dropped the backing table to the target keyspace, leaving a non-existent reference that resulted in errors creating Pull Requests. The mitigation was executing a command to drop the vschema reference to the dropped table, allowing Pull Request creation to resume.&lt;br /&gt;&lt;br /&gt;We are adding stronger pre-flight validation to our tooling to prevent similar issues and expanding lower-environment support to provide better test coverage end-to-end before promoting them to production. We&#039;re also fixing our backfill migration tooling to protect from this specific codepath..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065k2ouj0rmgk49b8jfc</id>
  <published>2026-07-24T16:17:31.492+00:00</published>
  <updated>2026-07-24T16:17:31.584+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065k2ouj0rmgk49b8jfc"/>
  <title>Disruption with some GitHub services</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 days, 7 hours and 19 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests, Third Party: GitHub → Pull Requests, Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:17:31</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for API Requests and Issues.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:19:39</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:22:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating timeouts to some GitHub services.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:28:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:20:10</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:26:47</var> GMT+0</small><br /><strong>Investigating</strong> -
  Copilot is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:27:51</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:40:27</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have applied a mitigation and are monitoring for recovery.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>16:41:59</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>17:16:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>17:24:27</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are seeing recovery across all services.</p>
<p><small>Jul <var data-var='date'> 24</var>, <var data-var='time'>17:36:50</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ. &lt;br /&gt;&lt;br /&gt;Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors: &lt;br /&gt;&lt;br /&gt;- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts. &lt;br /&gt;- 27% of GitHub issues interactions saw slow requests or timeouts. &lt;br /&gt;- 4% of GitHub Copilot requests experienced errors, though most automatically retry. &lt;br /&gt;- 4% of git push operations saw impacts during the affected window. &lt;br /&gt;- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were &lt; 1% in all cases.  &lt;br /&gt;&lt;br /&gt;We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36. &lt;br /&gt;&lt;br /&gt;This incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t065j2oug0rmg09hd9q8g</id>
  <published>2026-07-23T07:53:59.347+00:00</published>
  <updated>2026-07-23T07:53:59.492+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t065j2oug0rmg09hd9q8g"/>
  <title>Latency issues across a number of services</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 4 days, 9 hours and 20 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Webhooks, Third Party: GitHub → Pull Requests, Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>07:53:59</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded availability for Actions, Issues and Webhooks.</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>08:25:16</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pull Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>08:34:23</var> GMT+0</small><br /><strong>Investigating</strong> -
  We&#039;re currently investigating latency across multiple services. This can show as Actions jobs taking longer to start, Issues search serving stale results, and other listed services being similarly impacted..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:18:37</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Issues has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:19:49</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:22:21</var> GMT+0</small><br /><strong>Investigating</strong> -
  We identified the source of latency affecting multiple services and applied a fix. Issues and Actions are recovering, and remaining affected services are seeing improvement as processing backlogs clear. We are actively monitoring recovery across all services..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:27:35</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:35:17</var> GMT+0</small><br /><strong>Investigating</strong> -
  Webhooks is operating normally..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:39:11</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 23</var>, <var data-var='time'>09:39:19</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 23, 2026, between 07:08 and 09:39 UTC, several services experienced delays: 8% of actions workflow runs experienced an average run start delay of 10 minutes, 5% of webhook deliveries exceeded SLO, and code scanning, repos, notifications, issues and pull requests experienced increased latency over the life of the incident.    &lt;br /&gt;&lt;br /&gt;The root cause of the incident was a node of our background job processing system which did not recover after entering scheduled host maintenance.  The incident was mitigated by identifying the problematic shard and restoring its correct state, after which queue backlogs drained and services recovered.  &lt;br /&gt;&lt;br /&gt;To speed mitigation, we have added monitors for nodes in this unhealthy state after maintenance operations.  To prevent future recurrence, we are adapting our lifecycle automation to verify host rejoin after a scheduled reboot..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t066y2ouu0rmgw8fhbgad</id>
  <published>2026-07-22T20:43:43.789+00:00</published>
  <updated>2026-07-22T20:43:43.871+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t066y2ouu0rmgw8fhbgad"/>
  <title>Disruption with actions hosted runners</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 days, 13 hours and 43 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 22</var>, <var data-var='time'>20:43:43</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Jul <var data-var='date'> 22</var>, <var data-var='time'>20:47:02</var> GMT+0</small><br /><strong>Investigating</strong> -
  Approximately 3% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A small portion of these runs may fail after extended delays. We have identified the cause and are working on a mitigation..</p>
<p><small>Jul <var data-var='date'> 22</var>, <var data-var='time'>22:01:51</var> GMT+0</small><br /><strong>Investigating</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 22</var>, <var data-var='time'>22:01:57</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 22</var>, <var data-var='time'>22:09:26</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 22, 2026, between 19:36 UTC and 22:04 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 15% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 1% failed to start.&lt;br /&gt;&lt;br /&gt;At 21:49 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0dfp2p670rmgtubcgard</id>
  <published>2026-07-19T23:34:03.450+00:00</published>
  <updated>2026-07-20T00:49:33.644+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0dfp2p670rmgtubcgard"/>
  <title>Incident with GitHub Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 12 days, 21 hours and 59 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests, Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>00:49:33</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions and Pages are experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>00:52:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 19</var>, <var data-var='time'>23:34:03</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Jul <var data-var='date'> 19</var>, <var data-var='time'>23:37:51</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating degraded availability for GitHub Actions on github.com and in GHEC DR stamps. New workflows may delay or fail to start, and ongoing runs may fail. We will provide more information as soon as we can..</p>
<p><small>Jul <var data-var='date'> 19</var>, <var data-var='time'>23:57:14</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have identified the cause of failures in GitHub Actions and are working to restore service..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>00:07:41</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>00:20:23</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>00:48:31</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>03:27:10</var> GMT+0</small><br /><strong>Investigating</strong> -
  We’re seeing recovery across all impacted Actions runners. The team is continuing to monitor for global recovery..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>03:34:48</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>01:11:39</var> GMT+0</small><br /><strong>Investigating</strong> -
  We continue to work on mitigative efforts to restore Actions workflow runners, and have observed that the extended downtime has started to cause knock-on effects to other services.&lt;br /&gt;&lt;br /&gt;A separate incident was opened before we understood they were related. We will continue to post updates on this incident..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>01:37:10</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is operating normally..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>02:20:04</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is operating normally..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>02:23:53</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>04:43:50</var> GMT+0</small><br /><strong>Monitoring</strong> -
  Actions has fully recovered, and we will continue to monitor the platform to ensure stability..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>03:03:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is operating normally..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>02:43:43</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues, API Requests, and Pages have recovered. We are continuing to work on restoring GitHub Actions jobs using self-hosted or larger-hosted runners..</p>
<p><small>Jul <var data-var='date'> 20</var>, <var data-var='time'>04:44:03</var> GMT+0</small><br /><strong>Resolved</strong> -
  Between July 19, 2026, at 23:05 UTC and July 20, 2026, at 03:55 UTC, Actions self-hosted and larger runners were unable to connect to GitHub. During this period, Actions jobs were delayed or failed when trying to acquire a runner. Jobs using standard and Mac hosted runners were not affected. Reconnection traffic from affected runners also increased load on GitHub APIs, resulting in 3-4 seconds of additional average request latency and elevated 5xx error rates. &lt;br /&gt;&lt;br /&gt;The incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration that disrupted runner connectivity. We restored service by rotating the affected certificate. Recovery began at 02:45 UTC. By 03:55 UTC, queued workflow backlog had been processed and workflow delay rates returned to normal.&lt;br /&gt;&lt;br /&gt;To prevent recurrence, we are strengthening certificate renewal automation, adding fallback expiry monitoring and alerting, and improving circuit-breaker protections during runner API disruptions to reduce the risk of cascading impact to other APIs..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0dfd2p5p0rmgcgvbuyav</id>
  <published>2026-07-16T22:51:13.718+00:00</published>
  <updated>2026-07-17T00:00:47.214+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0dfd2p5p0rmgcgvbuyav"/>
  <title>Degraded REST API Availability</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 days, 10 hours and 54 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → API Requests</p>
    <p><small>Jul <var data-var='date'> 17</var>, <var data-var='time'>00:00:47</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 16</var>, <var data-var='time'>22:58:03</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are aware of degraded REST API availability and are investigating.</p>
<p><small>Jul <var data-var='date'> 17</var>, <var data-var='time'>00:14:08</var> GMT+0</small><br /><strong>Resolved</strong> -
  From 22:21 UTC - 23:50 UTC on July 16, 2026, the REST API experienced significant degradation.  During this period, about 39% of REST API requests failed with HTTP 500 level responses, with the errors peaking at 44.3%. &lt;br /&gt;&lt;br /&gt;We identified the issue as an infrastructure change that wrongly marked the majority of API backends in a single region as unhealthy.  As a result, requests routed to those backends failed before reaching the application layer.  &lt;br /&gt;&lt;br /&gt;To prevent this from happening again, we&#039;re improving our systems to catch this kind of invalid configuration before it reaches production. We&#039;ll also audit the related systems to make them more resilient to future changes, and we&#039;re increasing our monitoring sensitivity so we&#039;re alerted to problems like this sooner..</p>
<p><small>Jul <var data-var='date'> 16</var>, <var data-var='time'>22:51:13</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of impacted performance for some GitHub services..</p>
<p><small>Jul <var data-var='date'> 16</var>, <var data-var='time'>22:58:23</var> GMT+0</small><br /><strong>Investigating</strong> -
  API Requests is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 16</var>, <var data-var='time'>23:29:04</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to investigate an issue causing approximately 35% of REST API requests to fail.  Based on our current understanding, requests are not consistently reaching the application layer, resulting in failed requests returning HTML responses instead of the expected API response format. We are actively investigating the issue and will provide another update as soon as more information is available..</p>
<p><small>Jul <var data-var='date'> 17</var>, <var data-var='time'>00:14:04</var> GMT+0</small><br /><strong>Monitoring</strong> -
  As of 23:46 UTC, the REST API service is reachable and responding to requests normally..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0dfh2p5y0rmggma831ku</id>
  <published>2026-07-14T17:38:28.539+00:00</published>
  <updated>2026-07-14T17:38:28.636+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0dfh2p5y0rmggma831ku"/>
  <title>Incident with Webhooks</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 23 hours and 28 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Webhooks</p>
    <p><small>Jul <var data-var='date'> 14</var>, <var data-var='time'>17:38:28</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Webhooks.</p>
<p><small>Jul <var data-var='date'> 14</var>, <var data-var='time'>17:55:50</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Webhooks has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 14</var>, <var data-var='time'>17:55:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  Between 15:17 and 15:27 UTC, an ongoing deployment had an unintended side effect where webhook delivery states were not persisted in all cases, even when deliveries were accepted and processed. Customers may notice missing webhook deliveries in the UI, and these deliveries will not be retryable.&lt;br /&gt;&lt;br /&gt;Impact resolved automatically once the deployment completed..</p>
<p><small>Jul <var data-var='date'> 14</var>, <var data-var='time'>18:01:57</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 14, 2026, between 15:17 and 15:37 UTC, a rollout to GitHub&#039;s internal webhook delivery pipeline caused a subset of webhook delivery records to not be written to our webhook deliveries store after being processed and delivered successfully. Affected deliveries would be missing from the webhook delivery UI and API and won’t be available for redelivery.

The root cause was an uncoordinated rollout: a change to how delivery records are handed off between pipeline components was deployed before the upstream components producing those records were updated to match. While the rollout was in progress, affected records were silently skipped rather than persisted, with no automatic retry. The impact ended as soon as the rollout was completed.

About 2.4M delivery records were skipped (approximately 4% of the 20-minute impact window, 0.04% of a typical 24-hour period). Importantly, 95% of these skipped deliveries reached customer endpoints successfully, only the record of the delivery is missing. Of the ~5% that failed to reach customer endpoints, only ~1.4% (5,463) map to webhooks that retried their deliveries in the past 28 days.

To prevent recurrence, we are improving our automated detection of unsafe schema changes and tightening rollout coordination for changes that span multiple components in the pipeline..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0dg52p680rmg288dj96t</id>
  <published>2026-07-13T13:32:22.177+00:00</published>
  <updated>2026-07-13T13:32:22.327+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0dg52p680rmg288dj96t"/>
  <title>Actions runs are experiencing failures to start</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 21 hours and 34 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 13</var>, <var data-var='time'>13:32:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded availability for Actions.</p>
<p><small>Jul <var data-var='date'> 13</var>, <var data-var='time'>13:32:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 13</var>, <var data-var='time'>13:39:27</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 13</var>, <var data-var='time'>13:53:57</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 13, 2026, between 13:11 and 13:53 UTC, some customers experienced failures starting and running GitHub Actions workflows, which also affected Copilot cloud agent sessions and GitHub Pages builds since they depend on Actions. During the peak of the incident, 30% of Actions jobs failed to start and 2% were delayed more than 5 minutes. &lt;br /&gt;&lt;br /&gt;The incident was triggered by a configuration change in an internal autoscaling component that contained outdated capacity threshold values. This caused a critical Actions service to scale below its required baseline, reducing capacity for workflow processing. We identified the regression, rolled back the change, and restored service capacity. New workflow executions recovered by 13:39 UTC. Full recovery was reached by 13:53 UTC after the queued backlog was drained. &lt;br /&gt;&lt;br /&gt;To prevent recurrence, we have added deployment guardrails to validate that autoscaling inputs are current and to detect drift between planned and live scaling state before autoscaling changes are applied..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0dfc2p5m0rmgo4unx9md</id>
  <published>2026-07-09T04:34:24.849+00:00</published>
  <updated>2026-07-09T13:16:36.183+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0dfc2p5m0rmgo4unx9md"/>
  <title>Delays starting Actions runs</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 23 days, 5 hours and 52 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>13:16:36</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>13:39:48</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>13:41:26</var> GMT+0</small><br /><strong>Monitoring</strong> -
  Actions, Pages builds, Copilot Cloud Agent, and Copilot Code review have all recovered and are mitigated.&lt;br /&gt;&lt;br /&gt;We are continuing to monitor to ensure full recovery, and investigating the health of the affected infrastructure..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>04:34:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions.</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>12:46:13</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>10:07:04</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>12:36:33</var> GMT+0</small><br /><strong>Investigating</strong> -
  Pages is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>06:01:43</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to work on a mitigation..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>10:15:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  Approximately 30% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A smaller percentage of those are exhausting retries and failing to start..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>04:51:43</var> GMT+0</small><br /><strong>Investigating</strong> -
  Approximately 5% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A small portion of these runs may fail after extended delays. We have identified the cause and are working on a mitigation..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>12:01:59</var> GMT+0</small><br /><strong>Investigating</strong> -
  Approximately 30% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A smaller percentage of those are exhausting retries and failing to start.&lt;br /&gt;This has caused some customers to exceed their hosted compute concurrency and experience increased impact.&lt;br /&gt;&lt;br /&gt;We are continuing to working on infrastructure mitigations. &lt;br /&gt;&lt;br /&gt;Next update in one hour..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>12:54:24</var> GMT+0</small><br /><strong>Investigating</strong> -
  We&#039;re seeing Actions and Pages recovery. &lt;br /&gt;&lt;br /&gt;For a period of approximate 20 minutes ~96% of GitHub Actions runs on GitHub-hosted runners were failing to start, but has now recovered and we are seeing jobs processing.&lt;br /&gt;&lt;br /&gt;GitHub pages builds were also failing during that period, but Pages are still accessible.&lt;br /&gt;&lt;br /&gt;We are continuing to monitor for full recovery..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>13:17:04</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are continuing to monitor slow recovery in Actions and Pages builds as the system works through the high volume of backlog.&lt;br /&gt;&lt;br /&gt;Customers may see a small rate of  API and job failures as the system is recovering.&lt;br /&gt;&lt;br /&gt;Copilot Cloud Agent and Copilot Code Review also failed to start for approximately 30 minutes during this incident, and we are monitoring recovery.&lt;br /&gt;&lt;br /&gt;Pages were accessible throughout the incident..</p>
<p><small>Jul <var data-var='date'> 9</var>, <var data-var='time'>13:52:16</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 9, 2026, between 03:29 UTC and 13:39 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 8% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 2% failed to start.&lt;br /&gt;&lt;br /&gt;At 13:39 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0kic2pb80rmglykkjle0</id>
  <published>2026-07-07T14:14:39.949+00:00</published>
  <updated>2026-07-07T15:06:56.331+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0kic2pb80rmglykkjle0"/>
  <title>Actions and Codespaces APIs experiencing partial failures</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 5 days, 2 hours and 37 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Actions</p>
    <p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>15:06:56</var> GMT+0</small><br /><strong>Investigating</strong> -
  Customers accessing Actions runners and Codespaces REST APIs continue to see 500 errors a percentage of the time. Retries may be successful.&lt;br /&gt;&lt;br /&gt;We continue to investigate the source of these errors..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>14:14:40</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions and Codespaces.</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>16:00:01</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have rolled out the mitigation and are seeing recovery..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>14:32:57</var> GMT+0</small><br /><strong>Investigating</strong> -
  Codespaces is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>14:14:52</var> GMT+0</small><br /><strong>Investigating</strong> -
  Customers accessing the Actions and Codespaces REST APIs may see 500 class errors a small percentage of the time.&lt;br /&gt;&lt;br /&gt;Actions runs in progress are continuing successfully..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>14:47:06</var> GMT+0</small><br /><strong>Investigating</strong> -
  Customers accessing the Actions and Codespaces REST APIs may see 500 class errors a percentage of the time.  Retries may be successful.&lt;br /&gt;&lt;br /&gt;Actions runs and codespaces  in progress are continuing successfully.&lt;br /&gt;&lt;br /&gt;We are continuing to investigate the source of these errors..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>15:50:02</var> GMT+0</small><br /><strong>Investigating</strong> -
  Customers will continue to see 500 errors for approximately 8% of Actions runner REST APIs and 13% of Codespaces REST APIs. Retries may be successful. We have identified a likely cause and are preparing a mitigation..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>16:02:02</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions and Codespaces has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jul <var data-var='date'> 7</var>, <var data-var='time'>16:17:16</var> GMT+0</small><br /><strong>Resolved</strong> -
  On July 7, 2026, between 14:01 UTC and 16:17 UTC the Actions and Codespaces REST APIs were degraded and returned intermittent 500-class errors for a percentage of requests. Error rates peaked at approximately 8% of Actions runner API requests and 13% of Codespaces API requests, though retries were frequently successful. In-progress Actions runs and Codespaces were not impacted and continued successfully. This was due to a recent change that did not deliver the expected performance and, under certain conditions, caused downstream errors.&lt;br /&gt;&lt;br /&gt;We mitigated the incident by rolling back the change, after which the affected services recovered.&lt;br /&gt;&lt;br /&gt;We are working to improve the resilience of our services to these conditions and to strengthen our monitoring to reduce our time to detection and mitigation of issues like this one in the future..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmt0t0kiq2pbc0rmgeswdgnrj</id>
  <published>2026-06-25T17:50:19.266+00:00</published>
  <updated>2026-06-25T17:50:19.621+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmt0t0kiq2pbc0rmgeswdgnrj"/>
  <title>Degradation with Webhooks, Pull Requests and Actions</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 13 hours and 32 minutes</p>
    <p><strong>Affected Components:</strong> Third Party: GitHub → Webhooks, Third Party: GitHub → Pull Requests, Third Party: GitHub → Actions</p>
    <p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>17:50:19</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are investigating reports of degraded performance for Actions, Pull Requests and Webhooks.</p>
<p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>18:27:23</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We identified an issue that caused degradation across multiple services including Webhooks, Pull Requests, Actions, and Issues. Customers may have experienced delays or failures with these services. We have applied mitigations and affected services have recovered..</p>
<p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>17:58:46</var> GMT+0</small><br /><strong>Investigating</strong> -
  Issues is experiencing degraded performance. We are continuing to investigate..</p>
<p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>17:53:09</var> GMT+0</small><br /><strong>Investigating</strong> -
  Actions is experiencing degraded availability. We are continuing to investigate..</p>
<p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>18:07:44</var> GMT+0</small><br /><strong>Monitoring</strong> -
  The degradation affecting Actions, Issues, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability..</p>
<p><small>Jun <var data-var='date'> 25</var>, <var data-var='time'>18:27:52</var> GMT+0</small><br /><strong>Resolved</strong> -
  On June 25, 2026, between 17:33 UTC and 17:55 UTC, our background job service experienced degradation which increased delays to pull requests, repository pushes, Actions workflows, and Webhooks, with delays peaking at 7m. The issue was caused by underlying hypervisor issues and an incoming traffic spike, causing service timeouts which led to a connection storm and continual rebalances. &lt;br /&gt;&lt;br /&gt;The issue was mitigated by replacing the problem node at 17:49, after which all services saw recovery by 18:07..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.starsling.dev,2005:Incident/cmqs9l62g01x32jprts10vv8c</id>
  <published>2026-06-24T13:38:00.000+00:00</published>
  <updated>2026-06-24T15:32:02.794+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.starsling.dev/incident/cmqs9l62g01x32jprts10vv8c"/>
  <title>Jobs are delayed from starting for specific customers.</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 8 hours and 20 minutes</p>
    <p><strong>Affected Components:</strong> GitHub Runners</p>
    <p><small>Jun <var data-var='date'> 24</var>, <var data-var='time'>15:32:02</var> GMT+0</small><br /><strong>Monitoring</strong> -
  We implemented a fix and are currently monitoring the result..</p>
<p><small>Jun <var data-var='date'> 24</var>, <var data-var='time'>13:38:00</var> GMT+0</small><br /><strong>Identified</strong> -
  We&#039;ve identified an issue with Github API rate limiting that causes jobs to be delayed up to 25mins for some customers. We are continuing to work on a fix for this incident..</p>
<p><small>Jun <var data-var='date'> 24</var>, <var data-var='time'>18:07:48</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved. Zero recurrence in the last 2 hours. .</p>
<p><small>Jun <var data-var='date'> 24</var>, <var data-var='time'>18:51:33</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have observed a reoccurrence and are currently investigating this incident..</p>
<p><small>Jun <var data-var='date'> 24</var>, <var data-var='time'>21:58:08</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved. Job queue times are nominal. We pinpointed the secondary factor as an issue with a specific cloud provider and shifted load..</p>

        ]]>
  </content>
</entry>

</feed>