Was dreaming my tool got hacked and was contacting a CC, I would try to find the reason and just get bothered by different issues that were not related.
I realized I had to wake up to see what was happening, and I would wake up in the dream, and inside the dream I would try to analyze what was happening. Weird problems would happen around my setup, and I really woke up. Came to the computer connected to the tool stare at it and realize it was a dream, went back to bed.
After a while I got the same dream happening again, different people same ambience, same issue. Woke up, 6h30 am, came to the computer and it came to me in thought “PROMPT INJECTION” the tool is vulnerable to prompt injection.
Claude missed this by the way. Even in the review.
I was over hours handling and trying to understand the logic of terraform plugin management to get to the following config for being able to use telmate/proxmox latest github source code, that supports vm_state = “running”
First you should configure a ~/.terraformrc
// Replace $USER by your username
disable_checkpoint = true
provider_installation {
filesystem_mirror {
path = "/home/$USER/.terraform.d/plugins"
include = ["registry.terraform.io/telmate/*"]
}
direct { # This will prevent terraform from checking the online registry, that it always insists on checking
exclude = ["registry.terraform.io/telmate/*"]
}
}
Full path for the plugin ( I was using something like version 2.9.15 you can use whatever you prefer.
What I have learned and how it is built. And why I’m moving away from this.
How it is built
My homelab started on a old laptop that did mostly e-mail server and http service. Eventually I decided to move from the current setup of the laptop and acquired a refurbished machine for testing purposes. An Old Dell R710, I already had a Tsunami P4 , 2GB RAM and a Laptop Insys with 4 GB of RAM only the Dell R710 has Proxmox the others have simple clean Debian and disks.
I will restrain from the specs because my main objective was seeing how it would run the infra-structure in nomad in the long term, If a benchmark was intended I would have used a public cloud solution and build it up from ansible and Terraform like my linode setup for elasticsearch.
The Install is rather simple nothing much was needed I added glusterfs and docker to the mix. An example of a final nomad.hcl can be seen here.
Mostly at the top I ran 8 services some except the http service were not to be scaled.
Some of the services that I ran has a failsafe environment.
MariaDB HTTP for my domain webtrees private docker registry Shinobi NVR system to use a bit more of the disks on the GlusterFS pie hole wiki
The system was replicated and established in a environment, like they were at the same Data Center but behind different networks (VPCs).
Above you can see the setup, basically there were 6 hosts(VMs) implemented inside proxmox a Virtual Bridge separates each /28 network from the rest for isolation and to mimic different locations or failsafe zones, I just wanted complexity increased. Wireguard tunnel interconnects all the different Virtual Bridge Interfaces implement in the router with using firewalls. Initial implementation would specify that 62.0/24 network was reachable only by itself. Later I added the Tsunami and Insys for “fun”.
MariaDB resisted well to the various power outages and overall GlusterFS moves. Shinobi also but the camera footage always seems to break somehow Will detail this in later shinobi related post. Where did all the system mostly failed? I started adding the old laptop with Debian only (4GB of RAM) the laptop has issues and loses connectivity, causing the cluster to lose a node, Tsunami makes a lot of noise and so I turn it off. The cluster mostly is stable If I have the two nodes defined in the bootstrap, at the beginning I would loose a lot of time recovering the state of the cluster but eventually this became easier has I learned.
nomad members join mnmonic
Since I’m lazy I used traefik to redirect the http, mariadb (from outside firewalled), DNS udp (testing pie-hole) to inside to the containers in the nomad cluster.
Nomad allowed me to specify to what node should a job go , so heavy stuff like pushing a Magento App to the R710 and making sure it wouldn’t try to load on the other hosts worked well.
I lost the initial nomad job for magento has others , that’s a reason for restarting this project.
Most of all I ran a lot of services and the most I did test was resiliency of the setup and my hability to recover it.
GlusterFS
Mostly one thing that failed was after I acquired a new Server, and moved from the laptop services to the server, the server move was safe the new server also uses proxmox, but you have to have the naming of every glusterfs node or the system won’t mount the one specific host properly.
Since I initially used terraform with ansible to build the infra in proxmox and cloud-init was the option for machine deployment and setup. *Takes 6 minutes to setup the infra in the Dell R710. I basically missed that the template from cloud-init was always reloaded or I missed something, and with that the new GlusterFS node would not be configured in the hosts file.
Oh! Let’s talk about DNS and consul.
DNS and Consul
As said above I used consul to deploy the systems and properly configure the service names and DNS for the machines to be reachable from the outside.
service {
name = "jenkins"
tags = [
"traefik.enable=true",
"traefik.protocol=http",
"traefik.http.routers.jenkins.rule=Host(jenkins.a1.xxx.net)",
"traefik.http.routers.jenkins.entrypoints=web",
"traefik.http.services.jenkins.loadbalancer.server.scheme=http",
"traefik.http.services.jenkins.loadbalancer.passhostheader=true",
"traefik.http.services.jenkins.loadbalancer.healthcheck.hostname=jenkins.a1.xxx.net",
"traefik.http.services.jenkins.loadbalancer.responseforwarding.flushinterval=10"
]
port = "http"
Above you can see a basic DNS configuration for jenkins, Yes , also run jenkins on the thing on top of glusterfs. Don’t run jenkins master in a cluster, run the agents! Not the Master! The Build Agents! Don’t run the docker-registry inside the cluster. Why? Same reason has DNS, first you need to resolve docker-registry, your containers will be always trying to hit the DNS Resolve of the docker-registry in the consul until docker-registry comes up.
On a nomad failure for several times I had to restart the DNS forwarder (that redirected the requests to consul) so it would retry the DNS, now thinking about it , I could have tried playing with the Cache TLS time on the forwarder. #TODO
The most trouble was after cluster rebuild where some containers would not register properly the new address, I have to play with this later now that I think about it. #TODO
All and All I had more issues due to name resolution, than to the system itself failing. #itsalwaysdns
Well I did not use variables hidden inside consul, that will be for a later test case.
Now namespaces and ACL’s I haven’t found a usage for that in my case except separating environments of production and staging, basically that’s where I’m going next.
The question is… To separate Jenkins or not? Basically I wan’t now to deploy this full containers infra-structure using jenkins, althought Gitlab is easier I see where Jenkins has potential it clearly beats rundeck with some features like having a finally statement in the jobs. Something that simple could have saved me some time ago.
Today I’m proposing a terraform plan for maintaining your infra-structure we will start by creating a debian image Operating System install, in a thin disk and then proceed to exporting such image to our computer that has the terraform client.
Some requirements
ESXi 7
debian iso uploaded to a folder inside ESXi (I used a folder named iso)
Activate ssh server service in ESXi 7
Create a new user account to manage your terraforms plan/apply process with System Administration access, this will prevent that in case that your account get’s locked out due to invalid tries in ssh you still can get access with the root account and restart the ssh service to return access to the account and maybe reset the password in case you forgot it.
Enabling ssh system in ESXi 7
Go to Host -> Actions -> Services -> Enable Secure Shell
Create user with System Administration role in ESXi
Go to Manage Security & Users , Add userPlace the user name , password and confirm the password.Go to Host , Actions , PermissionsPress Add userAdd user for host in the blank text box place the user name and in the combo box select the role Administrator or other.
Now that you have your terraform user. Upload your iso!
Uploading the ISO
Download an iso from debian.org
Right click on the Storage -> Select Browse Datastores
Upload to the Datastorage , you may create a folder isos so everything get’s a bit more organized.
Install the Debian that will become the template to export
select the datastore
The disk and memoryselect CD/DVD DRive 1 with Datastorage ISO file
choose your iso file location and select the iso file
If your ok with everything press next
Download and prepare the file for Terraform
Right click and export the image
Change the .ovf file to remove the nvram file and remove the ExtraConfig
place it in a ./local folder inside the terraform directory
Place the correct disk size for the image vmdk disk or ESXi will complain.
This somes up the first part.
Terraform implementation
For terraform I tested the josenk terraform provider
https://github.com/josenk/terraform-provider-esxi
Download terraform >= 0.13
https://github.com/odnanref/terraform_esxi/tree/master = base source code
Terraform breakdown
First you have a resource that references the provider esxi that is supplied by josenk module.
provider "esxi" {
alias = "eg_host1"
esxi_hostname = "192.168.1.6"
esxi_hostport = "22"
esxi_hostssl = "443"
esxi_username = "${var.eg_esxi1_username}"
esxi_password = "${var.eg_esxi1_password}"
}
Here you have two variables that refer to a first ESXi hypervisor and you add the ones that you need I added two for example.
Specifying the resource
Well a bad part is that terraform does not allow dynamic providers.
So you will need to create a resource per each different provider using an alias for each provider.
terraform {
required_version = “>= 0.13”
}
resource “esxi_guest” “vmtest” {
provider = esxi.eg_host1
guest_name = “eg1-test”
disk_store = “datastore2”
memsize = “8192”
numvcpus = 12
#
# Specify an existing guest to clone, an ovf source, or neither to build a bare-metal guest vm.
4m to create two hosts on a Gigabit network 🙂 SSD Disks
18s to destroy the enviroment
Importing from a existing vm on the ESXi
clone_from_vm = “debian” instead of using ovf_source
It runs two exports for each host created I don’t have to transfer the image over the network but this export of the existing image for a copy is a pain.
5m , only a minute more, but given that if we are on a remote to remote site on a bandwith contained enviroment this can be a life saver.
Summary
Terraform + esxi with this module can be a life saver on a multiple ESXi enviroment when you don’t have a vcenter thingy. Terraform in this case given it did not have a dynamic way to specify the providers. Anyway this can be easy solved with some ingenuity.
The tests were done on a HP Gen 8 Server with 48 GB RAM DDR3 one disk of 160 GB and a disk for the VMs SSD of 120 GB Kingston no special model.
The hardware was aquired by be for personal tests on Ebay on a UK refurbished systems vendor.
repadmin /showrepl to see state of sync between Servers
See diagnostics about current server Issues
dcdiag to see more errors about sync
Check Options related to syncing servers
repadmin /options <SERVER|DC01|DC02>
Don’t rollback Domain Controllers being them a VM or not they will lose sync if the other DC was updated before and may corrupt the database.
Restore from backup details:
Is the system used for backups only available access throught the dead DC? Point it’s DNS only to the available DC ( had some issues with Samba) Samba will freak and start overusing IO for network in broadcast when authentication is requested, more, samba will not start correctly when one DC is down change the resolv.conf
Backup took one hour and ~30 minutes to restore.
More the time it takes to copy over the network to the local machine.
Copy the Backup to the local machine before restoring!
Until this moment there is no clear standard way on upgrading the system for non-paying (supported) users.
Recommended approach is
create secondary machine with cloned image first
Make sure your new versions of the ELK stack are supported by Elastiflow
do a full apt update
full apt upgrade
When upgrading elastiflow be sure to read the documentation on how to update plugins.
after upgrading replace basic config files
Redo all steps of install to make sure nothing was missed ( AT THIS MOMENT THERE IS NO UPGRADE DOCUMENTATION 2020-09-21 )
RE-import the kibana template into kibana overwriting old objects definitions / Noticed that when I did some changes to logstash always had to re-import the template.
Issues found during upgrade
Cluster Health was “yellow” due to Elasticsearch trying to allocate data to replicas ( I haven’t got a replica ) Had to directly assign index with no_replicas
Deverá estar ligado para publicar um comentário.