mirror of
https://github.com/outbackdingo/patroni.git
synced 2026-08-25 14:53:37 +00:00
Further work on permanent physical slots (#2891)
- Fixed issues with has_permanent_slots() method. It didn't took into account the case of permanent physical slots for members, falsely concluding that there are no permanent slots. - Write to the status key only LSNs for permanent slots (not just for slots that exist on the primary). - Include pg_current_wal_flush_lsn() to slots feedback, so that slots on standby nodes could be advanced - Improved behave tests: - Verify that permanent slots are properly created on standby nodes - Verify that permanent slots are properly advanced, including DCS failsafe mode - Verify that only permanent slots are written to the `/status`
This commit is contained in:
@@ -11,7 +11,7 @@ Feature: dcs failsafe mode
|
||||
When I issue a GET request to http://127.0.0.1:8008/failsafe
|
||||
Then I receive a response code 200
|
||||
And I receive a response postgres0 http://127.0.0.1:8008/patroni
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"postgresql": {"parameters": {"wal_level": "logical"}}}
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"postgresql": {"parameters": {"wal_level": "logical"}},"slots":{"dcs_slot_1": null,"postgres0":null}}
|
||||
Then I receive a response code 200
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"slots": {"dcs_slot_0": {"type": "logical", "database": "postgres", "plugin": "test_decoding"}}}
|
||||
Then I receive a response code 200
|
||||
@@ -44,14 +44,18 @@ Feature: dcs failsafe mode
|
||||
@dcs-failsafe
|
||||
@slot-advance
|
||||
Scenario: check leader and replica are functioning while DCS is down
|
||||
Given logical slot dcs_slot_0 is in sync between postgres0 and postgres1 after 10 seconds
|
||||
Given I get all changes from physical slot dcs_slot_1 on postgres0
|
||||
Then physical slot dcs_slot_1 is in sync between postgres0 and postgres1 after 10 seconds
|
||||
And logical slot dcs_slot_0 is in sync between postgres0 and postgres1 after 10 seconds
|
||||
And DCS is down
|
||||
Then Response on GET http://127.0.0.1:8008/primary contains failsafe_mode_is_active after 12 seconds
|
||||
Then postgres0 role is the primary after 10 seconds
|
||||
And postgres1 role is the replica after 2 seconds
|
||||
And replication works from postgres0 to postgres1 after 10 seconds
|
||||
And I get all changes from logical slot dcs_slot_0 on postgres0
|
||||
And logical slot dcs_slot_0 is in sync between postgres0 and postgres1 after 20 seconds
|
||||
When I get all changes from logical slot dcs_slot_0 on postgres0
|
||||
And I get all changes from physical slot dcs_slot_1 on postgres0
|
||||
Then logical slot dcs_slot_0 is in sync between postgres0 and postgres1 after 20 seconds
|
||||
And physical slot dcs_slot_1 is in sync between postgres0 and postgres1 after 10 seconds
|
||||
|
||||
@dcs-failsafe
|
||||
Scenario: check primary is demoted when one replica is shut down and DCS is down
|
||||
@@ -70,15 +74,43 @@ Feature: dcs failsafe mode
|
||||
And postgres1 role is the primary after 25 seconds
|
||||
|
||||
@dcs-failsafe
|
||||
Scenario: check three-node cluster is functioning while DCS is down
|
||||
Scenario: scale to three-node cluster
|
||||
Given I start postgres0
|
||||
And I start postgres2
|
||||
Then "members/postgres2" key in DCS has state=running after 10 seconds
|
||||
And "members/postgres0" key in DCS has state=running after 20 seconds
|
||||
And Response on GET http://127.0.0.1:8008/failsafe contains postgres2 after 10 seconds
|
||||
And replication works from postgres1 to postgres0 after 10 seconds
|
||||
And replication works from postgres1 to postgres2 after 10 seconds
|
||||
|
||||
@dcs-failsafe
|
||||
@slot-advance
|
||||
Scenario: make sure permanent slots exist on replicas
|
||||
Given I issue a PATCH request to http://127.0.0.1:8009/config with {"slots":{"dcs_slot_0":null,"dcs_slot_2":{"type":"logical","database":"postgres","plugin":"test_decoding"}}}
|
||||
Then logical slot dcs_slot_2 is in sync between postgres1 and postgres0 after 20 seconds
|
||||
And logical slot dcs_slot_2 is in sync between postgres1 and postgres2 after 20 seconds
|
||||
When I get all changes from physical slot dcs_slot_1 on postgres1
|
||||
Then physical slot dcs_slot_1 is in sync between postgres1 and postgres0 after 10 seconds
|
||||
And physical slot dcs_slot_1 is in sync between postgres1 and postgres2 after 10 seconds
|
||||
And physical slot postgres0 is in sync between postgres1 and postgres2 after 10 seconds
|
||||
|
||||
@dcs-failsafe
|
||||
Scenario: check three-node cluster is functioning while DCS is down
|
||||
Given DCS is down
|
||||
Then Response on GET http://127.0.0.1:8008/primary contains failsafe_mode_is_active after 12 seconds
|
||||
Then Response on GET http://127.0.0.1:8009/primary contains failsafe_mode_is_active after 12 seconds
|
||||
Then postgres1 role is the primary after 10 seconds
|
||||
And postgres0 role is the replica after 2 seconds
|
||||
And postgres2 role is the replica after 2 seconds
|
||||
|
||||
@dcs-failsafe
|
||||
@slot-advance
|
||||
Scenario: check that permanent slots are in sync between nodes while DCS is down
|
||||
Given replication works from postgres1 to postgres0 after 10 seconds
|
||||
And replication works from postgres1 to postgres2 after 10 seconds
|
||||
When I get all changes from logical slot dcs_slot_2 on postgres1
|
||||
And I get all changes from physical slot dcs_slot_1 on postgres1
|
||||
Then logical slot dcs_slot_2 is in sync between postgres1 and postgres0 after 20 seconds
|
||||
And logical slot dcs_slot_2 is in sync between postgres1 and postgres2 after 20 seconds
|
||||
And physical slot dcs_slot_1 is in sync between postgres1 and postgres0 after 10 seconds
|
||||
And physical slot dcs_slot_1 is in sync between postgres1 and postgres2 after 10 seconds
|
||||
And physical slot postgres0 is in sync between postgres1 and postgres2 after 10 seconds
|
||||
|
||||
@@ -654,9 +654,10 @@ class KubernetesController(AbstractExternalDcsController):
|
||||
try:
|
||||
if group is not None:
|
||||
scope = '{0}-{1}'.format(scope, group)
|
||||
ep = scope + {'leader': '', 'history': '-config', 'initialize': '-config'}.get(key, '-' + key)
|
||||
rkey = 'leader' if key in ('status', 'failsafe') else key
|
||||
ep = scope + {'leader': '', 'history': '-config', 'initialize': '-config'}.get(rkey, '-' + rkey)
|
||||
e = self._api.read_namespaced_endpoints(ep, self._namespace)
|
||||
if key != 'sync':
|
||||
if key not in ('sync', 'status', 'failsafe'):
|
||||
return e.metadata.annotations[key]
|
||||
else:
|
||||
return json.dumps(e.metadata.annotations)
|
||||
|
||||
@@ -50,7 +50,7 @@ Feature: ignored slots
|
||||
And postgres1 has a logical replication slot named unmanaged_slot_1 with the test_decoding plugin after 2 seconds
|
||||
And postgres1 has a logical replication slot named unmanaged_slot_2 with the test_decoding plugin after 2 seconds
|
||||
And postgres1 has a logical replication slot named unmanaged_slot_3 with the test_decoding plugin after 2 seconds
|
||||
And postgres1 does not have a logical replication slot named dummy_slot
|
||||
And postgres1 does not have a replication slot named dummy_slot
|
||||
|
||||
# 3. After a failover the server (now a primary) still has the slot.
|
||||
When I shut down postgres0
|
||||
|
||||
@@ -3,32 +3,73 @@ Feature: permanent slots
|
||||
Given I start postgres0
|
||||
Then postgres0 is a leader after 10 seconds
|
||||
And there is a non empty initialize key in DCS after 15 seconds
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"slots": {"test_physical": {"type": "physical"}}, "postgresql": {"parameters": {"wal_level": "logical"}}}
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"slots":{"test_physical":0,"postgres0":0,"postgres1":0,"postgres3":0},"postgresql":{"parameters":{"wal_level":"logical"}}}
|
||||
Then I receive a response code 200
|
||||
And Response on GET http://127.0.0.1:8008/config contains slots after 10 seconds
|
||||
When I start postgres1
|
||||
And I start postgres2
|
||||
And I configure and start postgres3 with a tag replicatefrom postgres2
|
||||
Then postgres0 has a physical replication slot named test_physical after 10 seconds
|
||||
And I start postgres1
|
||||
And postgres0 has a physical replication slot named postgres1 after 10 seconds
|
||||
And postgres0 has a physical replication slot named postgres2 after 10 seconds
|
||||
And postgres2 has a physical replication slot named postgres3 after 10 seconds
|
||||
|
||||
@slot-advance
|
||||
Scenario: check that logical permanent slots are created
|
||||
Given I run patronictl.py restart batman postgres0 --force
|
||||
And I issue a PATCH request to http://127.0.0.1:8008/config with {"slots": {"test_logical": {"type": "logical", "database": "postgres", "plugin": "test_decoding"}}}
|
||||
And I issue a PATCH request to http://127.0.0.1:8008/config with {"slots":{"test_logical":{"type":"logical","database":"postgres","plugin":"test_decoding"}}}
|
||||
Then postgres0 has a logical replication slot named test_logical with the test_decoding plugin after 10 seconds
|
||||
|
||||
@slot-advance
|
||||
Scenario: check that permanent slots are created on the replica
|
||||
Scenario: check that permanent slots are created on replicas
|
||||
Given postgres1 has a logical replication slot named test_logical with the test_decoding plugin after 10 seconds
|
||||
Then Logical slot test_logical is in sync between postgres0 and postgres1 after 10 seconds
|
||||
And Logical slot test_logical is in sync between postgres0 and postgres2 after 10 seconds
|
||||
And Logical slot test_logical is in sync between postgres0 and postgres3 after 10 seconds
|
||||
And postgres1 has a physical replication slot named test_physical after 2 seconds
|
||||
And postgres2 has a physical replication slot named test_physical after 2 seconds
|
||||
And postgres3 has a physical replication slot named test_physical after 2 seconds
|
||||
|
||||
@slot-advance
|
||||
Scenario: check that permanent slots are advanced on the replica
|
||||
Scenario: check permanent physical slots that match with member names
|
||||
Given postgres0 has a physical replication slot named postgres3 after 2 seconds
|
||||
And postgres1 has a physical replication slot named postgres0 after 2 seconds
|
||||
And postgres1 has a physical replication slot named postgres3 after 2 seconds
|
||||
And postgres2 has a physical replication slot named postgres0 after 2 seconds
|
||||
And postgres2 has a physical replication slot named postgres3 after 2 seconds
|
||||
And postgres2 has a physical replication slot named postgres1 after 2 seconds
|
||||
And postgres1 does not have a replication slot named postgres2
|
||||
And postgres3 does not have a replication slot named postgres2
|
||||
|
||||
@slot-advance
|
||||
Scenario: check that permanent slots are advanced on replicas
|
||||
Given I add the table replicate_me to postgres0
|
||||
And I get all changes from physical slot test_physical on postgres0
|
||||
When I get all changes from logical slot test_logical on postgres0
|
||||
And I get all changes from physical slot test_physical on postgres0
|
||||
Then Logical slot test_logical is in sync between postgres0 and postgres1 after 10 seconds
|
||||
And Physical slot test_physical is in sync between postgres0 and postgres1 after 10 seconds
|
||||
And Logical slot test_logical is in sync between postgres0 and postgres2 after 10 seconds
|
||||
And Physical slot test_physical is in sync between postgres0 and postgres2 after 10 seconds
|
||||
And Logical slot test_logical is in sync between postgres0 and postgres3 after 10 seconds
|
||||
And Physical slot test_physical is in sync between postgres0 and postgres3 after 10 seconds
|
||||
And Physical slot postgres1 is in sync between postgres0 and postgres2 after 10 seconds
|
||||
And Physical slot postgres3 is in sync between postgres2 and postgres0 after 20 seconds
|
||||
And Physical slot postgres3 is in sync between postgres2 and postgres1 after 10 seconds
|
||||
And postgres1 does not have a replication slot named postgres2
|
||||
And postgres3 does not have a replication slot named postgres2
|
||||
|
||||
@slot-advance
|
||||
Scenario: check that only permanent slots are written to the /status key
|
||||
Given "status" key in DCS has test_physical in slots
|
||||
And "status" key in DCS has postgres0 in slots
|
||||
And "status" key in DCS has postgres1 in slots
|
||||
And "status" key in DCS does not have postgres2 in slots
|
||||
And "status" key in DCS has postgres3 in slots
|
||||
|
||||
Scenario: check permanent physical replication slot after failover
|
||||
Given I shut down postgres0
|
||||
Given I shut down postgres3
|
||||
And I shut down postgres2
|
||||
And I shut down postgres0
|
||||
Then postgres1 has a physical replication slot named test_physical after 10 seconds
|
||||
And postgres1 has a physical replication slot named postgres0 after 10 seconds
|
||||
And postgres1 has a physical replication slot named postgres3 after 10 seconds
|
||||
|
||||
@@ -51,7 +51,7 @@ Feature: standby cluster
|
||||
When I issue a GET request to http://127.0.0.1:8010/patroni
|
||||
Then I receive a response code 200
|
||||
And I receive a response replication_state streaming
|
||||
And postgres1 does not have a logical replication slot named test_logical
|
||||
And postgres1 does not have a replication slot named test_logical
|
||||
|
||||
Scenario: check switchover
|
||||
Given I run patronictl.py switchover batman1 --force
|
||||
|
||||
@@ -110,7 +110,7 @@ def replication_works(context, primary, replica, time_limit):
|
||||
context.execute_steps(u"""
|
||||
When I add the table test_{0} to {1}
|
||||
Then table test_{0} is present on {2} after {3} seconds
|
||||
""".format(int(time()), primary, replica, time_limit))
|
||||
""".format(str(time()).replace('.', '_').replace(',', '_'), primary, replica, time_limit))
|
||||
|
||||
|
||||
@then('there is a "{message}" {level:w} in the {node} patroni log')
|
||||
|
||||
+16
-2
@@ -1,3 +1,4 @@
|
||||
import json
|
||||
import time
|
||||
|
||||
from behave import step, then
|
||||
@@ -36,8 +37,9 @@ def has_logical_replication_slot(context, pg_name, slot_name, plugin, time_limit
|
||||
assert False, f"Error looking for slot {slot_name} on {pg_name} with plugin {plugin}"
|
||||
|
||||
|
||||
@then('{pg_name:w} does not have a logical replication slot named {slot_name}')
|
||||
def does_not_have_logical_replication_slot(context, pg_name, slot_name):
|
||||
@step('{pg_name:w} does not have a replication slot named {slot_name:w}')
|
||||
@then('{pg_name:w} does not have a replication slot named {slot_name:w}')
|
||||
def does_not_have_replication_slot(context, pg_name, slot_name):
|
||||
try:
|
||||
row = context.pctl.query(pg_name, ("SELECT 1 FROM pg_replication_slots"
|
||||
" WHERE slot_name = '{0}'").format(slot_name)).fetchone()
|
||||
@@ -89,3 +91,15 @@ def has_physical_replication_slot(context, pg_name, slot_name, time_limit):
|
||||
pass
|
||||
time.sleep(1)
|
||||
assert False, f"Physical slot {slot_name} doesn't exist after {time_limit} seconds"
|
||||
|
||||
|
||||
@step('"{name}" key in DCS has {subkey:w} in {key:w}')
|
||||
def dcs_key_contains(context, name, subkey, key):
|
||||
response = json.loads(context.dcs_ctl.query(name))
|
||||
assert key in response and subkey in response[key], f"{name} key in DCS doesn't have {subkey} in {key}"
|
||||
|
||||
|
||||
@step('"{name}" key in DCS does not have {subkey:w} in {key:w}')
|
||||
def dcs_key_does_not_contain(context, name, subkey, key):
|
||||
response = json.loads(context.dcs_ctl.query(name))
|
||||
assert key not in response or subkey not in response[key], f"{name} key in DCS has {subkey} in {key}"
|
||||
|
||||
Reference in New Issue
Block a user