• kungen@feddit.nu
    link
    fedilink
    arrow-up
    3
    ·
    10 hours ago

    We don’t usually do it for fun, but I’m sure my teammate who has petitioned deploying a “chaosmonkey”-thingy to our prod clusters would love it 😁

    • flambonkscious@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      2
      ·
      9 hours ago

      I really wish we could get to that level of maturity (cries in healthcare)

      Not sure if we should take the approach of biting the bullet and forcing it on ourselves or slowly growing to it (under a hostile govt that sacked ~30% of the countries ICT team)

      • Peffse@lemmy.world
        link
        fedilink
        arrow-up
        1
        ·
        5 hours ago

        back when my team had proper coverage, we used to do chaosmonkey style tests for hosted health systems. We’d schedule an event with the hospital IT staff. Say… 1 hour on a specific day near midnight. The local IT staff would put out a notice that for the hour the system may be inaccessible. Then we’d start a test suite, as ungracefully as possible. We’d run a abort shutdown command on one server, see how RHCS would handle it. Bring it back up stable. Silo a database instance. See how that handled. Yank out the Red1 cable from the switch, see how the network balanced the new traffic over Red2/Blue. That kind of stuff. We caught a few problems that way.

        The clients hated it, but it really made the system more bulletproof for when a new recruit ran a command in the wrong window and took down an unrelated resource.

    • pivot_root@lemmy.world
      cake
      link
      fedilink
      arrow-up
      3
      ·
      edit-2
      10 hours ago

      Your teammate has the right approach here. People will quickly learn to stop relying on the dev environment in production deployments when the environment becomes unreliable.